Case Study · Independent Build
AI Restaurant Receptionist: the model proposes, the server decides
When a language model takes an order, something has to own the facts: the items, the prices, what was read back and what the caller agreed to. This is my independent build of a voice ordering agent for a demo restaurant: Vapi runs the conversation, a Rails server owns the order, and every claim below comes from the project's tests, logs, recorded browser calls and the first call to the deployed app, not from customers.
At a glance
- What it is
- A voice ordering agent for a demo restaurant. Vapi handles speech-to-text, the language model and text-to-speech; a Rails server owns every fact about the order, and the model's tool calls are requests.
- Where it runs
- Live at restaurant-receptionist.railsfanatics.com, deployed with Kamal to a shared server with its own Postgres. The homepage is public; the voice console is for signed-in operators, because every call is a paid voice session.
- How it was tested
- With real Vapi web calls started from a browser console, made by me, plus a frozen baseline, replays of recorded calls, a scripted evaluation and one call to the deployed app. No phone line, no SMS sent, no customers.
- Hardest problems
-
- What the agent said and what the order contained drifting apart, with nothing noticing.
- Making every tool call a validated, recorded request against a versioned cart.
- A model that submitted before the caller answered, in 2 of the 4 test calls that reached checkout, and again on the first production call.
- Showing that failure on demand, which a prompt did not do in two attempts.
- Stack
- Ruby on Rails 8.1, PostgreSQL, Hotwire (Turbo, Stimulus) and Action Cable, Devise; Vapi for speech-to-text, the language model (gpt-5-mini) and text-to-speech.
- My role
- Sole engineer: I designed, built and verified it, using AI coding agents (Claude) under my direction and review.
- More
- The live app · the repository · the August build log of the first prototype
The demo, in under three minutes
The problem
An early version could take an order by voice. In one recorded call the agent said "I'll add garlic knots" without calling the tool that adds them, read back a $16 order and submitted it. What the agent said and what the order contained had come apart, and nothing in the system noticed.
A caller cannot see the cart, and the model chooses when to call a tool and when to speak. So the question became: in a transaction run by a language model, what must the server decide?
The order belongs to the server
The agent's eight tools reach Rails through one authenticated webhook. Each call is validated, executed once per tool-call id, and recorded in the same transaction as the business change; the model gets a result or a structured error it can speak, never exception text.
- Prices come from the menu, never from tool arguments.
- Every cart change increments a cart version.
- The read-back is written by the server. Reading the cart back returns a sentence the server wrote, and records which version was read.
- A submit is accepted only at that version. A submitted order cannot be changed by the voice tools.
The menu lookup fell from 51 database queries and 4,554 bytes to 3 queries and 1,526 bytes, and from 17.7 to 0.9 ms (median of 50 runs, development database). Median server-side time per tool is 2 to 6 ms; voice round-trip latency was not measured.
Watching the model and the server disagree
A browser console starts a real Vapi web call and shows the live transcript, labelled not authoritative, beside the order board and the server's event stream, labelled authoritative, updated over Action Cable. A phone line and SMS were out of scope by plan: the console is the instrument, not a product screen.
What the live calls showed
Nine browser calls in Phase 1 (one an invalid run) turned up real, recurring failures. In 2 of the 4 calls that reached checkout the model submitted before the caller answered. It announced actions it never took, stalled after "One moment.", and spoke its reasoning aloud. Each became either a server rule or a documented limitation. In a later call it also added an item the caller had only asked about; the server recorded exactly what was requested, and the read-back stated it.
The server now refuses a submit unless the conversation history shows a caller turn after the
read-back; missing history fails closed. On a later recorded call the unchanged assistant read the order
back and submitted it in the same response. The server refused it with
customer_confirmation_required, the order stayed open at version 1, and after the caller's
yes the submit was accepted. The gate checks turn-taking, not what the caller said.
I first counted 3 premature submits out of 4. Re-checking each call against the voice platform's own model logs showed one was not premature: the model had the caller's yes, but the history stamped it 0.1 s after the submit. The count is 2 of 4, and that timing quirk is now a documented limitation of the gate.
In production
The app runs at restaurant-receptionist.railsfanatics.com, deployed with Kamal to the server that already hosts my other applications. It has its own Postgres container; the one service it shares is the TLS proxy, which the deploy did not restart, and the other applications kept running throughout. The homepage and sign-in are public and everything else needs a signed-in operator. The webhook needs its secret, and browsers reach a separate production Vapi assistant only through a public key restricted to this origin and that assistant.
Before the first deploy I booted the production image locally against a throwaway database. That rehearsal found a real defect: a thread setting that capped every database connection pool at 3, below the 5 that Solid Queue needs for its workers, so background jobs would never have started in production. It was fixed and pinned by a test before anything reached the server.
The first call to the deployed app was a plain order: one Margherita Pizza for pickup. Nine seconds after
get_cart returned the read-back, with no caller turn in between, the model called
submit_order. Rails refused it with customer_confirmation_required and kept the
order open at cart version 1. After the caller's yes the second submit was accepted, and the order was
confirmed at $14.00. The call lasted 144 seconds and cost $0.19; no background job failed, and no SMS was
queued (a web call has no phone number). One call is not a rate, but it is the same failure the test calls
showed, stopped by the same rule, on the deployed system.
A fault-injection test that didn't work
To show a refusal on demand I built a separate, labelled copy of the assistant whose prompt told it to submit straight after the read-back, on the same server and rules. In two live attempts it did not reproduce the failure: one call stalled before an order, and in the other the model waited for the caller. Meanwhile the normal assistant, told to wait, had done it on its own. A prompt changes how likely a mistake is; only the server's rules are certain.
How it is checked
- A frozen baseline. Before changing anything, 23 reliability rules were characterised against the original code and frozen at a git tag. They were rewritten to the intended behaviour, a 24th (the gate) was added, and the originals still pass against the tag.
- Real calls, replayed. Recorded calls are replayed through the webhook, and copies of real submit requests test the gate.
- A scripted evaluation. 15 conversations run through the real server path: 6 of 11 assistant claims were not reflected in the order, and both order invariants held in 15 of 15. It is scripted, not a model benchmark.
- Tests and CI. 472 automated tests with 0 failures, green in CI, including tests that pin the anonymous boundary and the production configuration; 93.5 % line coverage at the end of Phase 1, from 65.7 % at the baseline; all 20 Phase 1 acceptance criteria met. Fifteen live test calls cost $2.11.
What is shown and what is not
| Shown | Not shown |
|---|---|
| The server refusing a premature submit on a recorded live call, and on the first production call | How often the model makes that mistake |
| Both order invariants in 15 of 15 scripted conversations | A model evaluation or benchmark |
| Server-owned prices, cart versions, read-back, idempotent tool calls | A product with customers or a real restaurant |
| A deployed app behind sign-in, with its own database and a restricted production Vapi key | Load, uptime or a long production record |
| Real Vapi web calls from a browser console | A phone line, SMS delivery or call transfer |
| A fault-injection assistant on the same rules | That assistant failing on cue |
The server keeps the order equal to what was read back and answered; it cannot tell whether an add was what the caller meant. Conversational reliability (stalls, announced actions, misheard speech) is not established.
How it was built
Solo, from July to October 2026. I designed, built and verified it as a solo engineer, using AI coding agents (Claude) for much of the implementation, under my direction and review. I set the architecture, the reliability rules and the acceptance criteria, made every live call and decided what counts as evidence.
What this demonstrates for client work
If a model is going to change real data in your system (orders, bookings, payments, records), the work is in deciding which facts the model may never own, making every action a validated, recorded request, and testing against what the model actually does on live sessions, including the failures.
Related
- I'm Learning Voice AI by Building an AI Receptionist for a Restaurant, the August 2026 build log of the first prototype
- Keepford, a separate AI phone receptionist (Python, Retell, a live phone line, home services), about safety and booking decisions on a phone line; this build is about transactional state in an order
An independent build on a demo restaurant; calls are browser web calls made by the author, and figures are this project's tests, recorded calls and the first production call, not rates or load measurements. State as of 1 October 2026.
Facing a build like this?