Case Study · Independent Build

AI Restaurant Receptionist: the model proposes, the server decides

When a language model takes an order, something has to own the facts: the items, the prices, what was read back and what the caller agreed to. This is my independent build of a voice ordering agent for a demo restaurant: Vapi runs the conversation, a Rails server owns the order, and every claim below comes from the project's tests, logs, recorded browser calls and the first call to the deployed app, not from customers.

2 of 4 test calls reaching checkout where the model submitted before the caller answered
15 / 15 scripted conversations in which both order invariants held
51 → 3 database queries per menu lookup
472 automated tests, 0 failures

At a glance

What it is
A voice ordering agent for a demo restaurant. Vapi handles speech-to-text, the language model and text-to-speech; a Rails server owns every fact about the order, and the model's tool calls are requests.
Where it runs
Live at restaurant-receptionist.railsfanatics.com, deployed with Kamal to a shared server with its own Postgres. The homepage is public; the voice console is for signed-in operators, because every call is a paid voice session.
How it was tested
With real Vapi web calls started from a browser console, made by me, plus a frozen baseline, replays of recorded calls, a scripted evaluation and one call to the deployed app. No phone line, no SMS sent, no customers.
Hardest problems
  1. What the agent said and what the order contained drifting apart, with nothing noticing.
  2. Making every tool call a validated, recorded request against a versioned cart.
  3. A model that submitted before the caller answered, in 2 of the 4 test calls that reached checkout, and again on the first production call.
  4. Showing that failure on demand, which a prompt did not do in two attempts.
Stack
Ruby on Rails 8.1, PostgreSQL, Hotwire (Turbo, Stimulus) and Action Cable, Devise; Vapi for speech-to-text, the language model (gpt-5-mini) and text-to-speech.
My role
Sole engineer: I designed, built and verified it, using AI coding agents (Claude) under my direction and review.
More
The live app · the repository · the August build log of the first prototype

The demo, in under three minutes

AI Restaurant Receptionist demo (AI voice: elevenlabs.io). A clean recorded order, then the recorded refusal, then the fault-injection result. The clean call was an additional call made for footage, and the only clean one of the four normal-assistant calls in that phase. The narration is an AI-generated voice; captions are available.

The problem

An early version could take an order by voice. In one recorded call the agent said "I'll add garlic knots" without calling the tool that adds them, read back a $16 order and submitted it. What the agent said and what the order contained had come apart, and nothing in the system noticed.

A caller cannot see the cart, and the model chooses when to call a tool and when to speak. So the question became: in a transaction run by a language model, what must the server decide?

Who decides what. Vapi runs the conversation with speech-to-text, gpt-5-mini and text-to-speech; it proposes tool calls and hears back what actually happened. Each tool call is a request to Rails: Voice::ToolRunner takes one authenticated webhook, validates the arguments, is idempotent per tool-call id, records every execution as a ToolInvocation and returns structured errors, never exception text. OrderTaking, under row locks, takes prices from the menu, never from tool arguments, makes a new cart version on every change, writes the read-back and records its version, and accepts a submit only at that version after a caller turn. PostgreSQL holds orders, order items, tool invocations and call logs. A live browser console shows the Vapi transcript, marked not authoritative, beside the order board and server events from Rails.
Vapi runs the conversation; Rails owns every business fact. Calls are web calls from the browser console. Full diagram.

The order belongs to the server

The agent's eight tools reach Rails through one authenticated webhook. Each call is validated, executed once per tool-call id, and recorded in the same transaction as the business change; the model gets a result or a structured error it can speak, never exception text.

  • Prices come from the menu, never from tool arguments.
  • Every cart change increments a cart version.
  • The read-back is written by the server. Reading the cart back returns a sentence the server wrote, and records which version was read.
  • A submit is accepted only at that version. A submitted order cannot be changed by the voice tools.

The menu lookup fell from 51 database queries and 4,554 bytes to 3 queries and 1,526 bytes, and from 17.7 to 0.9 ms (median of 50 runs, development database). Median server-side time per tool is 2 to 6 ms; voice round-trip latency was not measured.

Watching the model and the server disagree

A browser console starts a real Vapi web call and shows the live transcript, labelled not authoritative, beside the order board and the server's event stream, labelled authoritative, updated over Action Cable. A phone line and SMS were out of scope by plan: the console is the instrument, not a product screen.

What the live calls showed

Nine browser calls in Phase 1 (one an invalid run) turned up real, recurring failures. In 2 of the 4 calls that reached checkout the model submitted before the caller answered. It announced actions it never took, stalled after "One moment.", and spoke its reasoning aloud. Each became either a server rule or a documented limitation. In a later call it also added an item the caller had only asked about; the server recorded exactly what was requested, and the read-back stated it.

The server now refuses a submit unless the conversation history shows a caller turn after the read-back; missing history fails closed. On a later recorded call the unchanged assistant read the order back and submitted it in the same response. The server refused it with customer_confirmation_required, the order stayed open at version 1, and after the caller's yes the submit was accepted. The gate checks turn-taking, not what the caller said.

A premature submit, refused, in one recorded call with the normal assistant. The caller says pickup and the model calls get_cart. Rails writes the read-back from the cart: one Margherita Pizza with extra cheese, total sixteen dollars, recorded as read at cart version 1. The model sends the read-back and submit_order in one response, with no caller turn after the read-back. Rails refuses with customer_confirmation_required: 0 caller turns since the last get_cart; nothing is submitted and the order stays open at version 1. The caller says yes, the model submits again, and Rails accepts: confirmed at version 1, with 1 caller turn after the read-back. Verified against Vapi's own model-request log.
One recorded call: the read-back and the submit in one model response, the refusal, then the caller's yes and the accepted submit. Full diagram.

I first counted 3 premature submits out of 4. Re-checking each call against the voice platform's own model logs showed one was not premature: the model had the caller's yes, but the history stamped it 0.1 s after the submit. The count is 2 of 4, and that timing quirk is now a documented limitation of the gate.

In production

The app runs at restaurant-receptionist.railsfanatics.com, deployed with Kamal to the server that already hosts my other applications. It has its own Postgres container; the one service it shares is the TLS proxy, which the deploy did not restart, and the other applications kept running throughout. The homepage and sign-in are public and everything else needs a signed-in operator. The webhook needs its secret, and browsers reach a separate production Vapi assistant only through a public key restricted to this origin and that assistant.

Before the first deploy I booted the production image locally against a throwaway database. That rehearsal found a real defect: a thread setting that capped every database connection pool at 3, below the 5 that Solid Queue needs for its workers, so background jobs would never have started in production. It was fixed and pinned by a test before anything reached the server.

The first call to the deployed app was a plain order: one Margherita Pizza for pickup. Nine seconds after get_cart returned the read-back, with no caller turn in between, the model called submit_order. Rails refused it with customer_confirmation_required and kept the order open at cart version 1. After the caller's yes the second submit was accepted, and the order was confirmed at $14.00. The call lasted 144 seconds and cost $0.19; no background job failed, and no SMS was queued (a web call has no phone number). One call is not a rate, but it is the same failure the test calls showed, stopped by the same rule, on the deployed system.

A fault-injection test that didn't work

To show a refusal on demand I built a separate, labelled copy of the assistant whose prompt told it to submit straight after the read-back, on the same server and rules. In two live attempts it did not reproduce the failure: one call stalled before an order, and in the other the model waited for the caller. Meanwhile the normal assistant, told to wait, had done it on its own. A prompt changes how likely a mistake is; only the server's rules are certain.

How it is checked

  • A frozen baseline. Before changing anything, 23 reliability rules were characterised against the original code and frozen at a git tag. They were rewritten to the intended behaviour, a 24th (the gate) was added, and the originals still pass against the tag.
  • Real calls, replayed. Recorded calls are replayed through the webhook, and copies of real submit requests test the gate.
  • A scripted evaluation. 15 conversations run through the real server path: 6 of 11 assistant claims were not reflected in the order, and both order invariants held in 15 of 15. It is scripted, not a model benchmark.
  • Tests and CI. 472 automated tests with 0 failures, green in CI, including tests that pin the anonymous boundary and the production configuration; 93.5 % line coverage at the end of Phase 1, from 65.7 % at the baseline; all 20 Phase 1 acceptance criteria met. Fifteen live test calls cost $2.11.

What is shown and what is not

ShownNot shown
The server refusing a premature submit on a recorded live call, and on the first production callHow often the model makes that mistake
Both order invariants in 15 of 15 scripted conversationsA model evaluation or benchmark
Server-owned prices, cart versions, read-back, idempotent tool callsA product with customers or a real restaurant
A deployed app behind sign-in, with its own database and a restricted production Vapi keyLoad, uptime or a long production record
Real Vapi web calls from a browser consoleA phone line, SMS delivery or call transfer
A fault-injection assistant on the same rulesThat assistant failing on cue

The server keeps the order equal to what was read back and answered; it cannot tell whether an add was what the caller meant. Conversational reliability (stalls, announced actions, misheard speech) is not established.

How it was built

Solo, from July to October 2026. I designed, built and verified it as a solo engineer, using AI coding agents (Claude) for much of the implementation, under my direction and review. I set the architecture, the reliability rules and the acceptance criteria, made every live call and decided what counts as evidence.

What this demonstrates for client work

If a model is going to change real data in your system (orders, bookings, payments, records), the work is in deciding which facts the model may never own, making every action a validated, recorded request, and testing against what the model actually does on live sessions, including the failures.

Related

An independent build on a demo restaurant; calls are browser web calls made by the author, and figures are this project's tests, recorded calls and the first production call, not rates or load measurements. State as of 1 October 2026.

Facing a build like this?

← All case studies