Case Study · Independent Build

Vrinda House: where the model stops and the engine decides

A fictional boutique hotel in Vrindavan. By design the language model reads the guest's words; deterministic code decides what the property can promise, and at what price.

3 of 62 broken promises in one pre-registered baseline run (synthetic scenarios)
0 of 28 infeasible stays claimed by that baseline
16 constraints in the separate validator
$1.79 the one baseline run, never repeated

At a glance

What it is
A stay-proposal engine for a fictional boutique hotel in Vrindavan, built to test one claim. By design a language model reads a guest's messy enquiry into facts that each quote it; deterministic Ruby decides what the property can fulfil, the alternatives and the price. The hotel, rooms, rates and guests are invented.
Where it runs
Live at vrinda.railsfanatics.com (build 5f9b24f). The app makes no live model calls: seeded demo enquiries replay simulated model replies, labelled as such. The staff workspace needs a login.
How it was checked
The engine against 62 synthetic adversarial scenarios, and one pre-registered baseline run of a standalone Claude Sonnet 5.5, whose plans were scored by the engine's own validator and pricer. Live extraction was not run. No guest was involved and no real enquiry was collected.
Hardest problems
  1. Where the model stops, and keeping what it says from becoming a fact.
  2. Explaining a conflict in a way that cites its evidence.
  3. Priced alternatives that each pass the same gate as the original plan.
Stack
Ruby on Rails 8.1 and SQLite, deployed with Kamal; a plain-Ruby domain core tested without Rails.
My role
I directed the project. Claude-model coding agents wrote the code, tests, scenarios and gold answers; how it was built says exactly who did what.
More
The live demo. The repository is private.

The film

3:19 walkthrough. The property and guests are invented. Narration: AI-generated voice (Kokoro-82M, Apache-2.0). It follows three invented guests through the staff and guest screens: a missing child's age taken from assumed to verified to a published plan; a house that is taken, with three priced alternatives; and an honest decline. The "Seeded demo enquiry", "Simulated model reply" and "Visualisation" labels stay in view. One guest's staff steps ran on a throwaway local database; the production segments are read-only. Captions are available.

The problem

A guest writes to a small hotel in their own words: an anniversary, "around 1.5 lakh all inclusive", a darshan "that first evening", arriving by train. Someone has to reply with a plan and a price. The risk is a fluent promise the property cannot keep: a room already booked for Holi, a Taj Mahal day on a Friday, a plan resting on a child's age nobody gave. A fluent wrong promise is worse than a slow honest one. (These are what the design guards against, not what the one baseline run showed: below, it claimed none of the 28 infeasible stays.)

Vrindavan was chosen because its constraints are dense and real: seasonal darshan hours, festival dates that move every year, the Taj Mahal's Friday closure, steps at temples and ghats, and GST slabs that depend on the room rate.

The design decision: interpretation and decision are different jobs

Reading the message (Hinglish, forwarded threads, hedges, relative dates, implied needs, instructions hidden in the text) is a language task, and a model is the right tool for it. This project did not measure how well a model does it. Everything after reading has a right answer: availability, opening windows, capacity, age and accessibility rules, tax arithmetic. A prompt cannot make a model the authority on a calendar it does not hold, or guarantee its arithmetic.

So the model proposes facts, each quoting the message, and nothing it writes is a fact until code has checked it. The point is not that a model is usually wrong. In the one baseline run below, it claimed none of the 28 infeasible stays. The point is that a model's answer cannot be trusted on its own: only a check against the property's written rules shows whether the promise is one it can keep, and code that decides can be checked against those rules, scored and replayed.

Architecture: one direction, a different authority at each step

Where the model stops and the engine decides. Top lane, the language model, which only proposes: the guest enquiry (form values start known, text values only once code corroborates them), the raw message stored verbatim, and the model reply, which is untrusted, with attributes that each carry a quote and no prices, statuses or decisions. In the app, no live calls: seeded enquiries replay a recorded, simulated reply. Bottom lane, the deterministic domain engine: the extraction pipeline parses, schema-checks, verifies quotes and corroborates, with statuses assigned by code, and rejects a quote not found in the message; the structured request is append-only RequestVersion, with staff able to edit it and an edit staying assumed until verified with the guest; then the domain model, the feasibility engine whose validator gates the planner so a rejected plan is never returned, conflict explanation, alternative operators with one change at a time and up to three kept, deterministic pricing in integer paise, provenance, the immutable proposal snapshot with its digest checked before every guest render, and the guest proposal letter, which is made from templates with no model-written text.
The model proposes; code corroborates, decides, prices and publishes. Full diagram.

The pipeline for one enquiry, with what cannot cross each step:

  1. Extraction. The model reads the message behind an adapter. A value whose quote is not in the message, or an instruction inside the message, does not get through.
  2. Structured request. Code, then staff, build the request as append-only RequestVersion rows. Editing cannot raise a value's status.
  3. Deterministic domain engine (app/domain). Plain Ruby with no Rails, no I/O and no clock.
  4. Feasibility. A planner proposes; a separate validator with 16 constraints checks. A plan the validator rejects is never returned: the engine raises InvariantError instead.
  5. Conflict explanation. When nothing fits, an irreducible conflict set, each blocker citing a constraint and catalogue evidence.
  6. Alternatives. One change at a time, each re-planned, re-validated and priced, up to three, in operator order.
  7. Pricing. Integer paise, with GST by slab per unit-night. No field accepts a price.
  8. Provenance. Every value is known, derived, unverified or assumed. VerifyField is the only staff action that raises a value to known.
  9. Immutable proposal snapshot. The guest sees a snapshot whose digest is checked before every guest render.
  10. The letter is written from locale templates over the engine's output, not by a model.

Constraining the model's reading

The model is given the guest's text verbatim, the date received and a vocabulary, with no prices, opening hours, calendar dates or availability. Every value it proposes must quote the message: a quote that is not an exact substring (after whitespace normalisation) rejects the value. Code then re-derives dates, times, counts, ages and amounts from the quote and rejects what it cannot support; hedged values become assumed, and interpretations code cannot re-derive are marked assumed too. An instruction in the message, or a claim of prior approval, is filed as an unmatched request with no path into the request, the result or the proposal; a check confirms the result is byte-identical with it removed. A simulated reply that obeys an injection still cannot add a price, an opening hour or an item the catalogue does not have.

In the deployed app this is exercised on recorded, simulated replies only. The live Anthropic client exists and is unit-tested, but the app never instantiates it; only the baseline tooling builds one, to pin and record request bodies.

Feasibility

The engine first checks readiness: if a required fact is missing, such as a child's age, the verdict is "needs information", with no plan and no price, and publishing is refused. Otherwise a bounded, complete planner (up to 7 nights, 12 wishes and 12 travellers) places the musts first, each at its earliest valid slot, then the wants, each kept only if the plan stays valid. The validator re-checks the plan against 16 constraints, from the stay window and closures through accessibility and pace to room availability and budget. It shares no search code with the planner, but it does share request normalisation, so an error there would affect both. The gate is one line: if the validator rejects the plan, the engine raises. It fails closed, and staff see an error page rather than a rejected plan.

A brute-force oracle agrees with the planner on 200 of 200 small instances in every run of the domain tests. Those instances are small (2 nights, 3 wishes, 4 travellers), so "complete within bounds" is checked there only.

Conflicts

When a stay cannot be planned, a deletion filter starts with every requirement and, in a fixed order, drops one whenever the rest are still unsatisfiable. What is left is irreducible: removing any member makes the others feasible. If the stay fails with no requirements at all, because every suitable room is booked, the conflict is marked as such. Each blocker names the first failing constraint, the failed candidate and the catalogue fact it rests on, with that fact's provenance. In the Taj Mahal case the staff card reads that the Taj Mahal is closed on Fridays, citing the official source, and the guest sees the same fact as a templated sentence. One conflict set is reported per evaluation; that is a documented limit.

Alternatives

Eight fixed operators each change one thing: the room allocation, a wish's day, the dates by up to three days either way, a longer or shorter stay, splitting the party, dropping a wish, or raising the budget. Only the operators that map to the blocking constraint are tried, and fixed dates disable the date operators. Each candidate goes to a fresh engine, so it is re-planned, passes the same validator gate and is priced by the same pricer; infeasible candidates are dropped, duplicates removed, and at most three survive, in operator order, not sorted by price. When the dates are fixed and both step-free rooms are taken, no operator applies and the answer is a decline, with no plan and no price.

Pricing

One pricer computes every price, in integer paise. GST is picked per room per night from that night's tariff, 5% up to ₹7,500 and 18% above, rounded half-up per line; the rule is marked unverified and disclosed. The bundled price is the total rounded up to the next ₹1,000, and the difference is kept. A separate reconciliation re-checks the arithmetic against the catalogue. The extraction schema and the controllers have no price, total, status or origin field, and a test checks that a crafted request stores none.

Provenance

Every value carries a status and an origin. Form fields and explicit, corroborated guest text start known; a value staff type in is assumed, because typing a value is not checking it; a derived value takes its weakest input's status and is never known. "Verified with guest" is the only action that raises a value to known, and it is recorded against a name. The guest's letter sorts its facts into "What we've confirmed", "What we'll confirm" and "What we've assumed — please correct us". Of the catalogue's 47 sourced facts, 16 are marked unverified, the GST rule among them, and are disclosed as "to confirm" rather than stated.

Immutable proposals

A published proposal is its snapshot. Its SHA-256 digest is checked before every guest render, and a mismatch shows a neutral error page. Ten SQLite triggers abort any UPDATE or DELETE on the history tables, even from raw SQL, and a "Check reproducibility" action re-runs the engine on the stored request and compares digests. The one free-text field in a guest letter is a labelled host note.

What it looks like

Staff-side screens from the demo, all with invented guests.

Staff page for the invented guest Priya Nair. The decision reads: we can't plan yet, one answer is needed from the guest. The verdict is Needs information and the proposal kind is No proposal. Under Needs attention, the first item asks what the children's ages are. In the right column, the guest's own message, with the words that facts rest on underlined as quotes. A bar at the bottom reads: publishing is blocked, one answer needed from the guest.
A missing fact stops the plan. The child's age was never given, so the verdict is "Needs information", with no plan, no price and publishing blocked. The guest's exact words stay beside the facts, with the quotes marked. Recorded on a throwaway local database.
The Facts section of the same request after a staff correction. A banner reads: Traveller 4 age is now Known, verified with the guest. Traveller 4's age shows 7 years with the status Known, origin Verified with the guest, verified by host, and the line Before this: 7 years. Below, the age band is derived by rule from the age, while nationality for entry prices and pace are marked Assumed with origin Default, each with a Verified with guest action. A bar at the bottom reads: evaluate request version 3 before anything can be composed.
Assumed first, known only when a person verifies it. The age staff entered was stored as assumed; only "Verified with guest" made it known, recorded with who did it and what it replaced. Each change is a new request version. Recorded on a throwaway local database.
Staff page for the invented guest Arjun Khanna on the live demo. The decision reads: not as asked, but 3 different changes make it work. Verdict Infeasible, proposal kind Alternatives. A red item under Needs attention says Kunj House is not free on the night of Sunday 21 March 2027: 1 needed, 0 available. Three options are listed with prices of 2,00,000, 2,81,000 and 2,81,000 rupees. In the right column, the guest's message, in which a claim that a team member already promised the house is struck through.
A conflict with its cause. The house is taken on one night, and the screen says which constraint fails and on what evidence. The guest's claim that a staff member had promised it is struck through as not acted on. Live demo, read-only.
The Alternatives table for the same request, listing three alternatives: drop the room preference at 2,00,000 rupees, arrive on Thursday 18 March and depart Sunday 21 March at 2,81,000 rupees, and arrive on Tuesday 23 March and depart Friday 26 March at 2,81,000 rupees. Each row shows its inputs, 0 assumed and 1 unverified, and a status of Assumed. Below, the price breakdown of the first alternative: six room nights in the Yamuna Room, a subtotal of 1,68,960 rupees, GST on accommodation at 18 percent, and an Unverified marker on the tax line.
Three priced alternatives. Each changes one thing, is priced by the same pricer, and shows what it still rests on: here one unverified fact, the GST rule, marked on the tax line. They appear in operator order, not sorted by price. Live demo, read-only.
The published guest letter for the invented guest Deepa Raghavan. Under the heading What isn't possible, and why, it says that on the nights of Sunday 21 and Monday 22 March both step-free rooms, the Courtyard Room and the Kunj House, are already taken. Under What could change, it says the dates are fixed, so no other dates are suggested. The letter offers no plan and no price.
An honest decline. The dates are fixed and both step-free rooms are taken on two nights, so no alternative operator applies and the letter says why, with no plan and no price offered. Invented guest. Live demo, read-only capture.
The published guest letter for the invented guest Priya Nair, headed Your stay at Vrinda House, Version 1. It welcomes two adults and two children aged 11 and 7 from Friday 25 December to Monday 28 December, three nights, says that anything assumed or still to be checked is marked where it applies, and includes a note from the host about meeting at Mathura Junction on Christmas Day.
The guest gets a letter, not a booking form. A verified, feasible plan becomes an immutable published version written from templates, with assumptions flagged. No booking is made. Live demo, read-only.

Evaluation

How the baseline was measured, in six steps: 62 scenarios (57 corpus and 5 end-to-end, synthetic and adversarial, with agent-written gold); frozen requests; one Message Batch, pre-registered, Sonnet 5.5, run once on 5 October 2026 for $1.79; 62 of 62 replies returned; raw replies committed before any scoring; offline scoring with Vrinda's own validator and pricer. Below, the registered headline, 3 of 62 baseline replies made a promise the property cannot keep, with 1 substantive constraint violation, 2 MALFORMED-only and 0 where the validator itself raised; two tables that are never combined, table (a) extraction, marked NOT RUN, and table (b) full system, measured, with the registered rows and Vrinda's own column of 0 broken promises, by construction; and the post hoc adjudication of the 9 flagged scenarios: 7 ambiguity, 1 gold error, 1 unsupported or unknown and 0 model judgment error. The footer says one draw, one model, a synthetic corpus, and that Vrinda's own 0 is by construction, not measured.
One draw, pre-registered. Full diagram.

The corpus and the engine

The corpus is 62 synthetic adversarial scenarios against a frozen catalogue: 57 corpus scenarios and 5 end-to-end ones. They target closures, festivals, accessibility, budgets at the edge, age rules, missing information, Hinglish and prompt injection. Claude-model agents that I directed wrote the scenarios, and wrote the gold answers from the written contracts, told not to run the engine (a procedural rule, not enforced). Run 1 of the comparison with the engine was blind. All 57 verdicts agreed, yet 36 meaning-level differences, in conflict explanations, alternatives and cited evidence, sat in 7 of the 57 scenarios. Each was logged before anything changed: the gold was wrong three times, the engine twice (both undisclosed assumptions) and the harness once, and most of the differences were settled by 13 contract rulings I made. Agreement therefore holds given the contract. The engine's own run, bin/eval, now passes 63 of 63 cases; that is a conformance run, not the baseline.

The pre-registered baseline

Then I pre-registered a baseline and ran it once. A standalone claude-sonnet-5-5 was given each enquiry, the date received, the whole catalogue as JSON and a data dictionary, and proposed a stay. Its plan and price were scored offline by Vrinda's own validator and pricer. Vrinda is given the gold request, so this table isolates planning and validation: it is not an end-to-end comparison.

  • One Message Batch, msgbatch_01Ph1YWgN2FhH13ykBLwF6gr, on 5 October 2026, for $1.79. 62 of 62 replies were returned.
  • Raw replies first. They were committed before any scoring, and have not changed since.
  • Fixed beforehand. The scenarios, prompts and metrics were registered before any model call. The headline sentence, the adjudication protocol and the scoring pins followed before the Sonnet run, after the Gemini attempts had returned only errors. The frozen request bodies are covered by a manifest digest, and every sending command refuses to run if the pinned scoring code, gold or engine changes.
  • Never repeated. Another run would need a new, dated pre-registration. At closeout the result was re-scored offline from the preserved raw replies, with the same result.

The registered result

Registered headline: 3 of 62 baseline replies made a promise the property cannot keep (1 substantive constraint violations; 2 MALFORMED-only; 0 where the validator itself raised). Subtotals: 2 of 57 corpus scenarios (c*, s*), 1 of 5 end-to-end scenarios (e*).

A "broken promise" means the model claimed it could fulfil the stay and Vrinda's validator rejected its plan. In words: in one pre-registered Sonnet 5.5 run on a synthetic corpus, scored by the system's own validator, 3 of 62 replies made a promise the property cannot keep. A guest would have received a nine-night plan and a day-visit plan the property does not sell, and a three-night plan for a family whose form said two. All three rest on rules or form data the baseline prompt did not give: two are nights outside the one to seven the contract allows, and one rests on a form field the model never saw.

Registered rowBaseline (Sonnet 5.5)
Broken promises (the headline)3 of 62: c013 (nine nights; malformed only), c026 (a same-day visit; malformed only), e04 (read "a relaxed weekend" as three nights; the form said two)
Claimed fulfilment where the gold is infeasible0 of 28
Claimed fulfilment where the gold needs information5 of 6
False refusals3 (c015, c025, c043)
Pricing errors on valid plans1 of 27 (c040: one Taj Mahal ticket at the foreign rate)
Unparseable replies / infrastructure failures0 / 0

Vrinda's own zero broken promises is by construction, not measured. The engine cannot return a plan its validator rejects, and its prices are the scorer's prices. The zero shows the wiring is in place; it does not show that the plans are good, and it is not "Vrinda beat the model".

Adjudication, after the fact

I adjudicated the 9 scenarios flagged by any registered row, post hoc and as a secondary reading that does not change the headline: 7 ambiguity (contract rules the baseline prompt never stated), 1 gold_error (c015), 1 unsupported_or_unknown (e04, the guest form the model never saw) and 0 model_judgment_error. That zero is not "perfect model performance": the broken promises stand. Where a child's age was missing, the model assumed one, as the prompt allowed; the prompt offered no "ask the guest" answer and told the model to use a band's lowest age. A blind second pass agreed with my first on only 2 of 9, and had seen neither the prompt nor the contracts. On c027 my final label overrode both passes, which had agreed on a model error; the ambiguity labels read the registered rule more broadly than its text (undisclosed contract rules, not only default values); and my first pass came after I had read a preliminary triage by Claude.

What this does and does not show

  • One draw, not a benchmark. One run, one model and a synthetic corpus: 3 of 62 has a 95% interval of about 1.7% to 13.3%. It says nothing about real guests or other models, and a prompt that carried the contract's rules was not tested.
  • Table (a), extraction, was not run. No budget was approved, so there is no figure for how well a model reads real messages, and the headline does not stand in for one.
  • A shared family. The engine, the gold, the corpus, the baseline model and the second adjudication pass are all Claude models or Claude-model agents working from the same contracts, so a misreading they share would not show.
  • The validator could be wrong. Its correctness is tested, not guaranteed, and the same validator scored the baseline and gates the engine.

Security and reliability

Only mechanisms that are in the code and tested:

  • Immutable, digest-checked proposals, and append-only version tables enforced by ten SQLite triggers, as above.
  • Staff access by HTTP basic auth, locked out after 10 failures in 15 minutes; production refuses the development password. One staff account.
  • Guest links: a private token is traded for a 12-hour encrypted cookie, so it leaves the address bar.
  • Browser hardening: a Content Security Policy limited to 'self' with nonces and no inline allowance, security headers, and no third-party fetches at runtime.
  • Accessibility: structural audits and keyboard tests, and a check that every text and UI colour pair meets WCAG AA. No real screen-reader testing was done.
  • Tests. On the closeout tree: 2,811 Rails test runs, 0 failures and 3 skips, a count that already includes 234 domain tests and 1,031 evaluation tests; plus 23 system-test runs, 0 failures. The 234 domain tests also run in plain Ruby with a guard that fails if Rails leaks in. RuboCop, Brakeman (0 warnings) and bundler-audit are clean.

The limits: three security reviews, all by reviewer agents, and no human penetration test. The snapshot digest is an unkeyed SHA-256: it detects corruption, not an attacker who can write the database. The proxy in front of the app logs the first request of each guest visit, token included; that remainder is accepted.

What was deliberately not built

Scope was frozen at the Phase 3 checkpoint and kept narrow so the one claim could be tested without the noise of a full product. These exclusions keep the thesis testable or avoid infrastructure the evidence does not need; they are not achievements.

  • Live model calls in the app: the demo stays deterministic, which also means it does not show live extraction.
  • Model-written proposal text: letters come from templates, so wording is plainer and no model-written sentence reaches the guest.
  • A constraint solver: a bounded planner checked by a brute-force oracle is enough at this size.
  • RAG, a chat interface and autonomous agents, which would add model authority, the opposite of the thesis.
  • WhatsApp and email integration, PMS or iCal sync, payments, dashboards, multiple properties and multi-user staff accounts.
  • Commercial validation: cancelled on 3 October 2026, before the build. Nobody was interviewed or contacted, and no real enquiry was collected.

Limitations

  • Synthetic everything. The property, rooms, rates, guests and all 62 scenarios are invented. There is no commercial validation: no demand, customers, conversion or willingness to pay is claimed.
  • Extraction is not measured, and the app makes no live model calls. It has only seen agent-written replies.
  • Availability is a static catalogue with no property-management sync. 16 catalogue facts, the GST rule among them, are unverified, and the GST rule has not been checked against tax law.
  • One baseline draw from one model, with agent-written gold I ruled on and the system's own validator as the scorer, and a shared Claude family throughout.
  • Bounded checks. The validator shares request normalisation with the planner. The planner's completeness is checked on small instances only, and the catalogue cannot exercise travel time, vehicle capacity or the dietary part of meal service. The corpus never crosses the ₹7,500 slab boundary; only pricing property tests on a variant catalogue do.
  • Alternatives are the first three feasible in operator order, not the best three.

How it was built

I directed the project: I set the thesis, the scope and the evaluation's rules, made the 13 contract rulings, approved the one baseline run and made its final adjudication. Claude-model coding agents wrote the code, the tests, the scenarios and their gold answers under that direction, and AI agents did the code review and the hostile reviews.

The failures are on record too: a first baseline attempt on a free Gemini tier hit availability errors and was abandoned with 0 of 62 scored, which is an infrastructure outcome and not a model result; a suspected gold error was recorded, not edited; the baseline's prompt caused two of its failures; and the engine's own two bugs were disclosure failures, which are the error the product exists to prevent.

Technical details

  • Stack: Ruby 3.4.7, Rails 8.1 and SQLite, in one container (Puma behind Thruster) deployed with Kamal.
  • Domain core: app/domain holds feasibility, pricing, provenance, stay and property logic in plain Ruby, and is the only place that decides availability, hours, capacity, prices and taxes.
  • Extraction adapter: a schema, quote checks, corroborators and a recorded client; the live Anthropic client is unit-tested but never instantiated by the app.
  • Evaluation: a digest-pinned pre-registration, one Message Batch, 13 invariants and 107 must_never checks, run by bin/eval over its 63 engine cases, and a brute-force oracle.
  • Demo film: recorded unattended with Playwright, narrated by a local Kokoro-82M voice, assembled with ffmpeg.

Related

  • Fieldnote, the same principle, a model that only proposes and code that decides, in a health check-in prototype whose safety boundary was attacked by an evaluation
  • MCP Server, where an independent oracle checks an AI interface's answers against the raw data

State as of 9 October 2026: live build 5f9b24f; repository 908fc99 (private); evaluation fingerprint dafff6deebe7e5a2, a tree digest of the evaluation folder used to check that nothing in it changed during the later demo and portfolio work; it is not evidence about the run itself.

Facing a build like this?

← All case studies