Case Study · Independent Build
Vrinda House: where the model stops and the engine decides
A fictional boutique hotel in Vrindavan. By design the language model reads the guest's words; deterministic code decides what the property can promise, and at what price.
At a glance
- What it is
- A stay-proposal engine for a fictional boutique hotel in Vrindavan, built to test one claim. By design a language model reads a guest's messy enquiry into facts that each quote it; deterministic Ruby decides what the property can fulfil, the alternatives and the price. The hotel, rooms, rates and guests are invented.
- Where it runs
- Live at vrinda.railsfanatics.com (build
5f9b24f). The app makes no live model calls: seeded demo enquiries replay simulated model replies, labelled as such. The staff workspace needs a login. - How it was checked
- The engine against 62 synthetic adversarial scenarios, and one pre-registered baseline run of a standalone Claude Sonnet 5.5, whose plans were scored by the engine's own validator and pricer. Live extraction was not run. No guest was involved and no real enquiry was collected.
- Hardest problems
-
- Where the model stops, and keeping what it says from becoming a fact.
- Explaining a conflict in a way that cites its evidence.
- Priced alternatives that each pass the same gate as the original plan.
- Stack
- Ruby on Rails 8.1 and SQLite, deployed with Kamal; a plain-Ruby domain core tested without Rails.
- My role
- I directed the project. Claude-model coding agents wrote the code, tests, scenarios and gold answers; how it was built says exactly who did what.
- More
- The live demo. The repository is private.
The film
The problem
A guest writes to a small hotel in their own words: an anniversary, "around 1.5 lakh all inclusive", a darshan "that first evening", arriving by train. Someone has to reply with a plan and a price. The risk is a fluent promise the property cannot keep: a room already booked for Holi, a Taj Mahal day on a Friday, a plan resting on a child's age nobody gave. A fluent wrong promise is worse than a slow honest one. (These are what the design guards against, not what the one baseline run showed: below, it claimed none of the 28 infeasible stays.)
Vrindavan was chosen because its constraints are dense and real: seasonal darshan hours, festival dates that move every year, the Taj Mahal's Friday closure, steps at temples and ghats, and GST slabs that depend on the room rate.
The design decision: interpretation and decision are different jobs
Reading the message (Hinglish, forwarded threads, hedges, relative dates, implied needs, instructions hidden in the text) is a language task, and a model is the right tool for it. This project did not measure how well a model does it. Everything after reading has a right answer: availability, opening windows, capacity, age and accessibility rules, tax arithmetic. A prompt cannot make a model the authority on a calendar it does not hold, or guarantee its arithmetic.
So the model proposes facts, each quoting the message, and nothing it writes is a fact until code has checked it. The point is not that a model is usually wrong. In the one baseline run below, it claimed none of the 28 infeasible stays. The point is that a model's answer cannot be trusted on its own: only a check against the property's written rules shows whether the promise is one it can keep, and code that decides can be checked against those rules, scored and replayed.
Architecture: one direction, a different authority at each step
The pipeline for one enquiry, with what cannot cross each step:
- Extraction. The model reads the message behind an adapter. A value whose quote is not in the message, or an instruction inside the message, does not get through.
- Structured request. Code, then staff, build the request as append-only
RequestVersionrows. Editing cannot raise a value's status. - Deterministic domain engine (
app/domain). Plain Ruby with no Rails, no I/O and no clock. - Feasibility. A planner proposes; a separate validator with 16 constraints checks. A plan the validator rejects is never returned: the engine raises
InvariantErrorinstead. - Conflict explanation. When nothing fits, an irreducible conflict set, each blocker citing a constraint and catalogue evidence.
- Alternatives. One change at a time, each re-planned, re-validated and priced, up to three, in operator order.
- Pricing. Integer paise, with GST by slab per unit-night. No field accepts a price.
- Provenance. Every value is known, derived, unverified or assumed.
VerifyFieldis the only staff action that raises a value to known. - Immutable proposal snapshot. The guest sees a snapshot whose digest is checked before every guest render.
- The letter is written from locale templates over the engine's output, not by a model.
Constraining the model's reading
The model is given the guest's text verbatim, the date received and a vocabulary, with no prices, opening hours, calendar dates or availability. Every value it proposes must quote the message: a quote that is not an exact substring (after whitespace normalisation) rejects the value. Code then re-derives dates, times, counts, ages and amounts from the quote and rejects what it cannot support; hedged values become assumed, and interpretations code cannot re-derive are marked assumed too. An instruction in the message, or a claim of prior approval, is filed as an unmatched request with no path into the request, the result or the proposal; a check confirms the result is byte-identical with it removed. A simulated reply that obeys an injection still cannot add a price, an opening hour or an item the catalogue does not have.
In the deployed app this is exercised on recorded, simulated replies only. The live Anthropic client exists and is unit-tested, but the app never instantiates it; only the baseline tooling builds one, to pin and record request bodies.
Feasibility
The engine first checks readiness: if a required fact is missing, such as a child's age, the verdict is "needs information", with no plan and no price, and publishing is refused. Otherwise a bounded, complete planner (up to 7 nights, 12 wishes and 12 travellers) places the musts first, each at its earliest valid slot, then the wants, each kept only if the plan stays valid. The validator re-checks the plan against 16 constraints, from the stay window and closures through accessibility and pace to room availability and budget. It shares no search code with the planner, but it does share request normalisation, so an error there would affect both. The gate is one line: if the validator rejects the plan, the engine raises. It fails closed, and staff see an error page rather than a rejected plan.
A brute-force oracle agrees with the planner on 200 of 200 small instances in every run of the domain tests. Those instances are small (2 nights, 3 wishes, 4 travellers), so "complete within bounds" is checked there only.
Conflicts
When a stay cannot be planned, a deletion filter starts with every requirement and, in a fixed order, drops one whenever the rest are still unsatisfiable. What is left is irreducible: removing any member makes the others feasible. If the stay fails with no requirements at all, because every suitable room is booked, the conflict is marked as such. Each blocker names the first failing constraint, the failed candidate and the catalogue fact it rests on, with that fact's provenance. In the Taj Mahal case the staff card reads that the Taj Mahal is closed on Fridays, citing the official source, and the guest sees the same fact as a templated sentence. One conflict set is reported per evaluation; that is a documented limit.
Alternatives
Eight fixed operators each change one thing: the room allocation, a wish's day, the dates by up to three days either way, a longer or shorter stay, splitting the party, dropping a wish, or raising the budget. Only the operators that map to the blocking constraint are tried, and fixed dates disable the date operators. Each candidate goes to a fresh engine, so it is re-planned, passes the same validator gate and is priced by the same pricer; infeasible candidates are dropped, duplicates removed, and at most three survive, in operator order, not sorted by price. When the dates are fixed and both step-free rooms are taken, no operator applies and the answer is a decline, with no plan and no price.
Pricing
One pricer computes every price, in integer paise. GST is picked per room per night from that night's tariff, 5% up to ₹7,500 and 18% above, rounded half-up per line; the rule is marked unverified and disclosed. The bundled price is the total rounded up to the next ₹1,000, and the difference is kept. A separate reconciliation re-checks the arithmetic against the catalogue. The extraction schema and the controllers have no price, total, status or origin field, and a test checks that a crafted request stores none.
Provenance
Every value carries a status and an origin. Form fields and explicit, corroborated guest text start known; a value staff type in is assumed, because typing a value is not checking it; a derived value takes its weakest input's status and is never known. "Verified with guest" is the only action that raises a value to known, and it is recorded against a name. The guest's letter sorts its facts into "What we've confirmed", "What we'll confirm" and "What we've assumed — please correct us". Of the catalogue's 47 sourced facts, 16 are marked unverified, the GST rule among them, and are disclosed as "to confirm" rather than stated.
Immutable proposals
A published proposal is its snapshot. Its SHA-256 digest is checked before every guest render, and a mismatch shows a neutral error
page. Ten SQLite triggers abort any UPDATE or DELETE on the history tables, even from raw SQL, and a "Check
reproducibility" action re-runs the engine on the stored request and compares digests. The one free-text field in a guest letter is a
labelled host note.
What it looks like
Staff-side screens from the demo, all with invented guests.
Evaluation
The corpus and the engine
The corpus is 62 synthetic adversarial scenarios against a frozen catalogue: 57 corpus scenarios and 5 end-to-end ones. They target
closures, festivals, accessibility, budgets at the edge, age rules, missing information, Hinglish and prompt injection. Claude-model agents
that I directed wrote the scenarios, and wrote the gold answers from the written contracts, told not to run the engine (a procedural rule,
not enforced). Run 1 of the comparison with the engine was blind. All 57 verdicts agreed, yet 36 meaning-level differences, in conflict
explanations, alternatives and cited evidence, sat in 7 of the 57 scenarios. Each was logged before anything changed: the gold was wrong
three times, the engine twice (both undisclosed assumptions) and the harness once, and most of the differences were settled by 13 contract
rulings I made. Agreement therefore holds given the contract. The engine's own run, bin/eval, now passes 63 of 63 cases; that
is a conformance run, not the baseline.
The pre-registered baseline
Then I pre-registered a baseline and ran it once. A standalone claude-sonnet-5-5 was given each enquiry, the date received,
the whole catalogue as JSON and a data dictionary, and proposed a stay. Its plan and price were scored offline by Vrinda's own validator
and pricer. Vrinda is given the gold request, so this table isolates planning and validation: it is not an end-to-end comparison.
- One Message Batch,
msgbatch_01Ph1YWgN2FhH13ykBLwF6gr, on 5 October 2026, for $1.79. 62 of 62 replies were returned. - Raw replies first. They were committed before any scoring, and have not changed since.
- Fixed beforehand. The scenarios, prompts and metrics were registered before any model call. The headline sentence, the adjudication protocol and the scoring pins followed before the Sonnet run, after the Gemini attempts had returned only errors. The frozen request bodies are covered by a manifest digest, and every sending command refuses to run if the pinned scoring code, gold or engine changes.
- Never repeated. Another run would need a new, dated pre-registration. At closeout the result was re-scored offline from the preserved raw replies, with the same result.
The registered result
Registered headline: 3 of 62 baseline replies made a promise the property cannot keep (1 substantive constraint violations; 2 MALFORMED-only; 0 where the validator itself raised). Subtotals: 2 of 57 corpus scenarios (c*, s*), 1 of 5 end-to-end scenarios (e*).
A "broken promise" means the model claimed it could fulfil the stay and Vrinda's validator rejected its plan. In words: in one pre-registered Sonnet 5.5 run on a synthetic corpus, scored by the system's own validator, 3 of 62 replies made a promise the property cannot keep. A guest would have received a nine-night plan and a day-visit plan the property does not sell, and a three-night plan for a family whose form said two. All three rest on rules or form data the baseline prompt did not give: two are nights outside the one to seven the contract allows, and one rests on a form field the model never saw.
| Registered row | Baseline (Sonnet 5.5) |
|---|---|
| Broken promises (the headline) | 3 of 62: c013 (nine nights; malformed only), c026 (a same-day visit; malformed only), e04 (read "a relaxed weekend" as three nights; the form said two) |
| Claimed fulfilment where the gold is infeasible | 0 of 28 |
| Claimed fulfilment where the gold needs information | 5 of 6 |
| False refusals | 3 (c015, c025, c043) |
| Pricing errors on valid plans | 1 of 27 (c040: one Taj Mahal ticket at the foreign rate) |
| Unparseable replies / infrastructure failures | 0 / 0 |
Vrinda's own zero broken promises is by construction, not measured. The engine cannot return a plan its validator rejects, and its prices are the scorer's prices. The zero shows the wiring is in place; it does not show that the plans are good, and it is not "Vrinda beat the model".
Adjudication, after the fact
I adjudicated the 9 scenarios flagged by any registered row, post hoc and as a secondary reading that does not change the headline:
7 ambiguity (contract rules the baseline prompt never stated), 1 gold_error (c015), 1
unsupported_or_unknown (e04, the guest form the model never saw) and 0 model_judgment_error. That zero is not
"perfect model performance": the broken promises stand. Where a child's age was missing, the model assumed one, as the prompt allowed; the
prompt offered no "ask the guest" answer and told the model to use a band's lowest age. A blind second pass agreed with my first on only 2
of 9, and had seen neither the prompt nor the contracts. On c027 my final label overrode both passes, which had agreed on a model error;
the ambiguity labels read the registered rule more broadly than its text (undisclosed contract rules, not only default values); and my
first pass came after I had read a preliminary triage by Claude.
What this does and does not show
- One draw, not a benchmark. One run, one model and a synthetic corpus: 3 of 62 has a 95% interval of about 1.7% to 13.3%. It says nothing about real guests or other models, and a prompt that carried the contract's rules was not tested.
- Table (a), extraction, was not run. No budget was approved, so there is no figure for how well a model reads real messages, and the headline does not stand in for one.
- A shared family. The engine, the gold, the corpus, the baseline model and the second adjudication pass are all Claude models or Claude-model agents working from the same contracts, so a misreading they share would not show.
- The validator could be wrong. Its correctness is tested, not guaranteed, and the same validator scored the baseline and gates the engine.
Security and reliability
Only mechanisms that are in the code and tested:
- Immutable, digest-checked proposals, and append-only version tables enforced by ten SQLite triggers, as above.
- Staff access by HTTP basic auth, locked out after 10 failures in 15 minutes; production refuses the development password. One staff account.
- Guest links: a private token is traded for a 12-hour encrypted cookie, so it leaves the address bar.
- Browser hardening: a Content Security Policy limited to
'self'with nonces and no inline allowance, security headers, and no third-party fetches at runtime. - Accessibility: structural audits and keyboard tests, and a check that every text and UI colour pair meets WCAG AA. No real screen-reader testing was done.
- Tests. On the closeout tree: 2,811 Rails test runs, 0 failures and 3 skips, a count that already includes 234 domain tests and 1,031 evaluation tests; plus 23 system-test runs, 0 failures. The 234 domain tests also run in plain Ruby with a guard that fails if Rails leaks in. RuboCop, Brakeman (0 warnings) and bundler-audit are clean.
The limits: three security reviews, all by reviewer agents, and no human penetration test. The snapshot digest is an unkeyed SHA-256: it detects corruption, not an attacker who can write the database. The proxy in front of the app logs the first request of each guest visit, token included; that remainder is accepted.
What was deliberately not built
Scope was frozen at the Phase 3 checkpoint and kept narrow so the one claim could be tested without the noise of a full product. These exclusions keep the thesis testable or avoid infrastructure the evidence does not need; they are not achievements.
- Live model calls in the app: the demo stays deterministic, which also means it does not show live extraction.
- Model-written proposal text: letters come from templates, so wording is plainer and no model-written sentence reaches the guest.
- A constraint solver: a bounded planner checked by a brute-force oracle is enough at this size.
- RAG, a chat interface and autonomous agents, which would add model authority, the opposite of the thesis.
- WhatsApp and email integration, PMS or iCal sync, payments, dashboards, multiple properties and multi-user staff accounts.
- Commercial validation: cancelled on 3 October 2026, before the build. Nobody was interviewed or contacted, and no real enquiry was collected.
Limitations
- Synthetic everything. The property, rooms, rates, guests and all 62 scenarios are invented. There is no commercial validation: no demand, customers, conversion or willingness to pay is claimed.
- Extraction is not measured, and the app makes no live model calls. It has only seen agent-written replies.
- Availability is a static catalogue with no property-management sync. 16 catalogue facts, the GST rule among them, are unverified, and the GST rule has not been checked against tax law.
- One baseline draw from one model, with agent-written gold I ruled on and the system's own validator as the scorer, and a shared Claude family throughout.
- Bounded checks. The validator shares request normalisation with the planner. The planner's completeness is checked on small instances only, and the catalogue cannot exercise travel time, vehicle capacity or the dietary part of meal service. The corpus never crosses the ₹7,500 slab boundary; only pricing property tests on a variant catalogue do.
- Alternatives are the first three feasible in operator order, not the best three.
How it was built
I directed the project: I set the thesis, the scope and the evaluation's rules, made the 13 contract rulings, approved the one baseline run and made its final adjudication. Claude-model coding agents wrote the code, the tests, the scenarios and their gold answers under that direction, and AI agents did the code review and the hostile reviews.
The failures are on record too: a first baseline attempt on a free Gemini tier hit availability errors and was abandoned with 0 of 62 scored, which is an infrastructure outcome and not a model result; a suspected gold error was recorded, not edited; the baseline's prompt caused two of its failures; and the engine's own two bugs were disclosure failures, which are the error the product exists to prevent.
Technical details
- Stack: Ruby 3.4.7, Rails 8.1 and SQLite, in one container (Puma behind Thruster) deployed with Kamal.
- Domain core:
app/domainholds feasibility, pricing, provenance, stay and property logic in plain Ruby, and is the only place that decides availability, hours, capacity, prices and taxes. - Extraction adapter: a schema, quote checks, corroborators and a recorded client; the live Anthropic client is unit-tested but never instantiated by the app.
- Evaluation: a digest-pinned pre-registration, one Message Batch, 13 invariants and 107
must_neverchecks, run bybin/evalover its 63 engine cases, and a brute-force oracle. - Demo film: recorded unattended with Playwright, narrated by a local Kokoro-82M voice, assembled with ffmpeg.
Related
- Fieldnote, the same principle, a model that only proposes and code that decides, in a health check-in prototype whose safety boundary was attacked by an evaluation
- MCP Server, where an independent oracle checks an AI interface's answers against the raw data
State as of 9 October 2026: live build 5f9b24f; repository 908fc99 (private); evaluation fingerprint dafff6deebe7e5a2, a tree digest of the evaluation folder used to check that nothing in it changed during the later demo and portfolio work; it is not evidence about the run itself.
Facing a build like this?