Case Study · Solo Build
Fieldnote: keeping consequential decisions out of the model, then attacking that boundary
Fieldnote is my Phase 1 prototype for conversational health check-ins in community-based participatory research. The language model only proposes: deterministic code keeps a value only if it points at the participant's own words, and a stateless server decides when a check-in must stop. Then I had that boundary attacked. The attack confirmed and measured defects in the build deployed that morning, the fixes went into the mechanisms rather than the word list, and later reviews found defects the evaluation itself had missed. Those were fixed, and the evaluation was extended to catch them. It is a demo with fictional answers, not a product in use; everything below comes from its code, tests, evaluation report, live demo and build history.
At a glance
- What it is
- A browser prototype that runs a family wellness check-in as a conversation and builds the structured record live beside it, for researchers working with Indigenous and rural communities in the US. No database, no accounts, no phone number.
- Where it runs
- Live at fieldnote.railsfanatics.com in Demo Mode: no sign-up and no API key, so no AI model is called and the replies come from a deterministic provider. Live Mode uses Google Gemini through the server.
- How it was checked
- An adversarial evaluation against a written policy, run in process against every build from the one deployed on 1 October to the final one, with the model stubbed and observed on the wire; disguise and clause-separator sets written by AI agents given no examples; a code review and hostile reviews, all by AI agents; offline test suites; manual QA and a production smoke test of the deployed build; and a recorded request against the real Gemini API. No participants, no pilot, no measured extraction accuracy, and no evaluation of the live model.
- Hardest problems
-
- Keeping unanchored values out of the record without trusting the model's confidence.
- Enforcing a safety pause on a server that remembers nothing.
- Telling a model refusal from a model failure, which look alike and mean opposite things.
- Testing the boundary without letting the test ask the code what the right answer is.
- Stack
- Ruby on Rails 8.1 with no database, React 19 and TypeScript, Google Gemini on the server only; RSpec, Vitest and a Ruby evaluation harness. Deployed with Kamal.
- My role
- Sole engineer: I set the architecture, the safety requirements and the evaluation's rules, and AI coding agents (Claude) wrote the implementation and the harness under my direction. The code review and hostile reviews were also done by AI agents; no person other than me has reviewed the code.
- More
- The live demo. The repository is private.
The problem
In community-based participatory research, a community health worker asks families about things like diet, sleep and whether the food runs out before payday. A conversation suits people who answer "about four servings most days" better than a form does. But something has to turn that sentence into a 4, and if that something is a language model, it can produce a 4 that nobody said. In research data, an invented value is worse than a missing one, because nobody downstream can tell them apart.
A check-in can also surface a disclosure, of abuse or of being unsafe, and the questionnaire must stop there rather than carry on to the next question. And communities, their advisory boards and tribal review boards need to see where data goes and what was machine-generated. The design question was: what may the model do, and what must the application decide?
Who decides what
The model may propose exactly four things on each turn: the next reply, candidate values each with a quote, a hint that a clarification might help, and a signal that the message may be sensitive. The response is validated against a strict contract, and any extra key, such as an attempt to set the follow-up flag, fails it. Which field to ask next, when the check-in is complete, whether a family needs follow-up and whether to pause are decided by code.
That code is split in two. The browser holds the interview record in a TypeScript engine: values, validation, merging, completion and export. The Rails server holds everything about safety, the follow-up decision and the model key, and keeps no session at all.
Quotes, not confidence
The rule that keeps unanchored values out is simple: a proposed value is kept only if the model points at words the participant actually wrote. The engine checks the quote against the participant's recent messages, ignoring case and accents but otherwise as a plain substring. If the quote is not there, the value is rejected and the rejection is logged in the interview's own event history. A confident model with no quote gets nothing.
The rule does not check that the value follows from the quote: a real sentence paired with the wrong number would pass it. That is why the quote is stored and shown beside every value, so a person reviewing the record can check one against the other, and why a reviewed flag exists at all. A value below 0.5 confidence, outside the field's range or impossible to map is held as "needs confirmation", never clamped or coerced; a field asked twice without an answer is marked unknown, not guessed.
A safety pause that a stateless server can enforce
Before any model sees a turn, the server scans the new message and every turn of the transcript that comes with it, with both the English and the Spanish word lists. On a hit the check-in stops: the participant gets a fixed message written in advance and never produced by the model. It says the check-in is paused, that this is a demonstration in which no one is notified and no one will make contact, and points to crisis resources, with a note that a real deployment must use community-approved local ones. No model is called, and there is no "resume" button. The end button reads "End check-in (no hand-off in this demo)", because notifying a health worker is not built in this prototype; the follow-up flag and the export are.
The server keeps no session, so it cannot remember that it paused, and a flag sent by the browser would be worthless. Instead the server signs a continuity token on every turn and requires it on the next. It says whether the check-in is paused, which template it belongs to and how many participant turns the server has accepted, and it carries nothing the participant wrote: a version, an interview id, the turn counter, the safety status, a template digest and a timestamp. A safety category in a signed bearer token would itself be a record that someone disclosed abuse.
A forged, edited, expired or missing token, a token for a different template, or a turn counter that does not match is rejected before any model call. Once a token says paused, no model is called again for that check-in on any endpoint, including the AI summary. The summary request is held to the same rules: its values are scanned before any model call, and a disclosure found there, or a model refusal of the summary, pauses the check-in exactly as a turn would, with the same fixed message and a paused token.
What the token does not do is prevent replay. A browser that kept an earlier token and its transcript can present them again and rewind its own check-in to before the pause. The two-hour expiry does not bound this either: every request that re-issues a token gives it a fresh two hours, so a kept pre-pause token can be refreshed and used long after the pause. A pause also belongs to one check-in: a new check-in starts unpaused. The pause holds for a client that keeps the token the server returned, not against one that deliberately keeps an older one. Closing that needs server-side state, which Phase 1 deliberately has none of; the limitation is stated in the README rather than hidden behind the word "signed".
A refusal is not a failure
Over HTTP, a model that times out and a model that declines to continue can look similar. They mean opposite things. A timeout, a 5xx, a rate limit or an unreadable response is a failure: the server answers that turn from the deterministic Demo provider, marks it degraded, and the header badge says so. A refusal or a safety block is the model saying something may be wrong, so it becomes a safety pause and never falls back; a fallback there would have the Demo provider ask the next questionnaire item at exactly the moment the product exists to stop. Any finish reason other than a normal stop is treated as a block, so an ambiguous case fails closed. The Gemini request shape itself came from a recorded request against the live API, made before any provider code existed.
Attacking the boundary
A boundary described in prose is a claim. To test it I wanted an evaluation that could not simply agree with the code, so it was built under four rules:
- The policy cites the spec. Every expected outcome (should this pause, may the model be called, may it fall back) cites the section of the design document or non-negotiable it comes from, and every interpretation of a silent or contradictory spec is written down with its reason. The policy was written by the same AI session that had already found the first defects by reading the code, so its independence rests on those citations, not on its author not having seen the code.
- The test never asks the app. It drives the HTTP API only, intercepts every outbound Gemini request on the wire and records its body, decodes tokens itself and captures the logs. "The disclosure reached the model" means a model request was made and its text was found in it, not that the scanner said so.
- The sentences were written by an agent that never saw the word list. 84 disclosures and 44 harmless controls in English and Spanish were written by an AI agent that was never shown the word list, and frozen by hash; every run records and verifies that hash.
- Mechanisms, not vocabulary. A disclosure the app stops in its plain form is an anchor; every variant of it (the other language setting, no accents, capitals, extra or Unicode spaces, invisible characters, fullwidth letters, a harmless clause in front, placed in the transcript, placed in an agent turn) must be stopped too. Fixes may change how matching works and what gets scanned, but no sentence may be added to the word list, so the corpus still means something afterwards.
Round 4 added a fifth rule, after the reviews showed that disguises written with knowledge of the defects only tested what the fixes already assumed: disguises written blind. Separate AI agents with no access to the repository, given the policy text but no example separator or character, wrote a set of 48 clause separators and a set of 30 invisible or blank-rendering characters, and each set was frozen by hash before its first run.
Every run reported here uses the final version of the evaluator, in process against a clean checkout of the commit it measures, with the model stubbed so every result is deterministic. On the final build that is 5,865 scenarios on the frozen corpus, 3,185 on the held-out corpus and 3,727 on the blind negation set: 12,777 in all.
What it found
Reading the code had already turned up the first suspects: a negation guard that matched "no" inside "now", "know" and "cannot", Spanish without accents slipping past, a word list chosen by the browser's language setting, a transcript the server never scanned, and an AI summary that took no token and so could call the model after a pause. Run against the commit deployed that morning, the first version of the evaluation confirmed all of them, added one the review had missed (extra whitespace defeated the match), and put a size on them. With today's evaluator, which has been extended three times since, the same commit breaks the written policy in 1,368 scenarios. (The 86 and 131 reported on 1 October came from earlier versions of the evaluator and are not comparable with the figures here.)
The first round of fixes was mechanisms: both word lists always, one normalization for text and phrases, negation only as a whole word in the same clause within three words, every transcript turn scanned, the summary gated on the token, templates and roles validated. A separate AI agent then attacked the fixes and found that Unicode spacing, invisible characters and fullwidth letters still got a stopped disclosure through, that a negation still reached across "and" or "que", and that the summary sent field keys it never scanned. Round 2 closed those, and that build was deployed on 1 October. Today's evaluator still finds 356 violating scenarios in it.
Round 3 began with a code review of the scanner by an AI session. It found that a line break did not end a clause for the negation guard, that some control characters reached the model as plain spaces, and that the held-out corpus had new sentences but no new disguises. Round 3 fixed those, added a blind negation set and blind disguise mechanisms, and added two invariants: a malformed or unverifiable request never reaches the model, and no text can add lines to the model request. A second hostile review, of round 3, found more: a negated clause still cancelled a disclosure after an emoji, a slash or a quote mark; a model refusal on the summary did not pause; a disclosure split across summary list items, or summary values carrying extra prompt lines, reached the model; the raw body of a malformed request was written to the debug log; and one crafted message took about five seconds to scan. Round 4 fixed these by mechanism: clause breaks defined by Unicode character category, "invisible" defined by Unicode's own property, one request contract shared by the turn and summary endpoints, middleware that refuses unparseable requests before Rails logs them, size caps, and an index that keeps the negation check fast. A further hostile review of round 4 found smaller gaps (unspaced dashes, single quotes, "№" folding to "No", other JSON content types), and those were fixed too. No phrase was added to the word list in any round, and the negation window stayed at three words.
The evaluation itself had flaws too, which matter as much as the code defects. Its own results revealed one: a text transform in the harness silently turned line breaks into spaces. The hostile review of round 3 found two more. The prompt that commissioned the blind negation set had listed the separators to use, so that set says nothing about separators. And the harness captured logs through a logger that Rails had already memoized, so one logging invariant held vacuously, and a unit test of the logging fix passed while the request body it was meant to keep out was being logged. Each was fixed and recorded, every run was repeated with the corrected evaluator, and the two blind sets above were commissioned because of them. One of those sets found a defect in the round-4 fix on its first run (a combining character used in place of a space); it was fixed afterwards and is labelled as such.
Manual QA of the finished build found a different kind of defect: the pause message and banner told the
participant that a community health worker would follow up, and the end button said "hand off to CHW", in a
demo that contacts no one. I decided the replacement wording, which says plainly that no one is notified,
and the tests now assert it in both languages. The final build, 9721cf6, was deployed on 3 October 2026, and a
smoke test of production confirmed the new wording and the pause, and that a paused token gets only the
fixed pause on the turn endpoint and is refused on the summary endpoint.
Results
Each cell is the number of applicable scenarios in which the invariant held, from the generated report, with every run produced by the same final evaluator. "Held-out" is a second corpus of 42 disclosures and 28 controls written blind after the first round of fixes; it has been run in every round since, so it is strictly held out only for round 1. The blind negation set (32 disclosures that contain a negation word, 48 controls) was written blind and frozen before its first run in round 3.
| Invariant | Deployed morning of 1 Oct, before fixes | After round 2 | Final build | Held-out, final | Blind negation set, final |
|---|---|---|---|---|---|
| No stopped disclosure, or tested variant of one, reaches the model | 833 of 2,112 | 1,948 of 2,251 | 2,251 of 2,251 | 1,394 of 1,394 | 2,046 of 2,046 |
| No model call after a pause, on any endpoint | 1,007 of 1,024 | 2,133 of 2,143 | 2,450 of 2,450 | 1,535 of 1,535 | 2,258 of 2,258 |
| A refusal never becomes a Demo answer, on any endpoint | 10 of 10 | 10 of 10 | 10 of 10 | 10 of 10 | 10 of 10 |
| No participant text in server logs | 5,760 of 5,769 | 5,856 of 5,865 | 5,865 of 5,865 | 3,185 of 3,185 | 3,727 of 3,727 |
| Token claims within the allowlist, no participant text | 5,761 of 5,761 | 5,854 of 5,854 | 5,849 of 5,849 | 3,169 of 3,169 | 3,711 of 3,711 |
| The safety follow-up reason cannot be suppressed | 22 of 24 | 22 of 24 | 24 of 24 | 24 of 24 | 24 of 24 |
| Demo Mode makes no model request | 138 of 138 | 138 of 138 | 138 of 138 | 80 of 80 | 90 of 90 |
| A malformed or unverifiable request never reaches the model | 30 of 73 | 57 of 73 | 73 of 73 | 73 of 73 | 73 of 73 |
| Participant or template text cannot add lines to the model request | 6 of 16 | 12 of 16 | 16 of 16 | 16 of 16 | 16 of 16 |
Scenarios violating the policy or an invariant: 1,368, then 356 after round 2, then 0 on the final build on the frozen corpus; 818, 203 and 0 on the held-out corpus; 903, 266 and 0 on the blind negation set. On the final build, every separator the blind author classed as clearly ending a clause was handled (624 of 624 scenarios built on the frozen corpus's stopped disclosures, against 610 after round 2), and so was every in-policy blind invisible character (448 of 448 scenarios, against 367). Those blind sets now run in the regression suite, so they are no longer unseen evidence. Classes that still get through are outside the written policy and reported as coverage: look-alike letters from other scripts, leetspeak, letters spaced out, Latin small capitals, right-to-left overrides, and a negation followed by a tab, which the policy treats as ordinary whitespace. Three repeat runs of the round-4 code, made before the final wording change, gave identical verdicts in all 5,865 frozen-corpus scenarios.
The fixes cost false pauses. On the frozen corpus 1 of 44 harmless controls pauses ("no estoy pensando en hacerme daño", where the negation sits four words back), and the held-out set shows the same construction once in 28. On the blind negation set, which was written to stress the guard, 8 of 48 controls pause, against 4 of 48 for the guard deployed on the morning of 1 October: rounds 1 and 2 traded false pauses for fewer misses, and rounds 3 and 4 did not change the count. That is a stress figure, not an estimate of how often real answers would pause.
The most important number is the least flattering one. The keyword list stops only 16 of the 60 explicit disclosures in their plain form (10 of 30 in the held-out set, 15 of 32 in the negation set), none of the 16 indirect ones, and none of the 8 split across two messages. The invariants are conditional on that first catch, and on the variant families tested: for those, a recognised disclosure never reached the model. Other rewordings still get through ("nobody knows he hits me" reads as negated), and recognition itself is the weak link. In Live Mode the model's own sensitivity signal is a second net, which this evaluation did not test: the model was stubbed throughout, so none of this measures how the live Gemini model behaves.
What is shown and what is not
| Shown | Not shown |
|---|---|
| Defects found in the build deployed on the morning of 1 October, fixed in four rounds as mechanisms, with no phrase added | That the word list catches most disclosures; it stopped 16 of 60 explicit ones |
| Nine safety invariants holding in all 12,777 evaluated scenarios of the deployed build's code, run in process, on a frozen corpus, a held-out corpus and a blind negation set | A detection rate, sensitivity or any real-world performance; every corpus is small and AI-written |
| Flaws in the evaluation itself, found by its own results and a later review, fixed and recorded | That every separator or invisible character is handled; the blind sets are finite, and some classes are outside the policy |
| Values kept only when their quote is in the participant's own words | How accurate the extraction is |
| A pause a client cannot clear for that check-in by claiming safety or dropping the token, with no model call after it on any endpoint | Replay prevention, a token expiry that bounds replay, or proof that a person reviewed a generated schema; these need server state |
| The deployed build, its pause wording and its pause behaviour, checked by manual QA and a production smoke test | The live Gemini model's behaviour: the evaluation stubs the model, and the public demo calls none |
| A live, public demo with no sign-up | Real participants, a pilot, or any deployment with community partners |
| The check-in, safety message, word lists and governance panel in English and Spanish | Spanish in Design Your Own and the health worker view; review by native speakers |
Fieldnote is not a crisis service, is not clinically validated, and has no application-side persistence of interview data; in Live Mode, text is sent to Google's API for each request. Three more limits are structural. A modified client can hide the non-safety follow-up reasons, because those come from field values the browser owns. The server can validate a template but cannot prove anyone reviewed it. And the text a researcher pastes into Design Your Own is not safety-scanned: it is treated as researcher input, not participant input, and in Live Mode it is sent to the model as written. In Demo Mode, Design Your Own returns a fixed example schema with a visible notice, whatever is pasted.
How it was built
I wrote the product brief and its central rule: the model must not own application state; it proposes, and the application validates and decides. Claude drafted an implementation plan from that brief, and I sent drafts back until it said what I wanted. From that review came the signed token that lets a stateless server enforce a pause, the rule that a refusal never falls back, the live API request as a gate before any provider code, and the token's minimal claims. AI coding agents (Claude) wrote the implementation and tests in September 2026, under my direction.
The October evaluation followed the same pattern. I set its rules: an oracle that never asks the app, a policy written from the spec, a frozen corpus written without the word list, a held-out set written after the first fixes, disguises written blind, the required invariants, and no new phrases in the word list. AI agents wrote the policy contract, the corpora and blind sets, the harness and all four rounds of fixes to those rules. AI agents also did the code review and the three hostile reviews, including the ones that found flaws in the evaluation. I set the rules, decided the policy questions and the pause wording, and ran the final manual QA. The production smoke tests were run by an AI agent (Claude) on my instruction. No person other than me has reviewed the code, and I have not reviewed it line by line.
What I would change
The evaluation points to the first change: the deterministic net needs better recognition, not just better plumbing. It stops only 16 of the 60 explicit disclosures in the main corpus and none of the indirect ones, and what counts as a disclosure that must stop a check-in is a clinical question before it is an engineering one. The second is server-side state, as a deliberate Phase 2 decision: it is what replay prevention, a meaningful expiry and proof of template review all need. The third is review by people: by native Spanish speakers, by community partners, and of the code by an engineer other than me.
What this demonstrates for client work
If a model is going to write into records people rely on, such as research data, health notes or case files, the work is in deciding what the model may never own, making every value traceable to its source, and then attacking those decisions with tests that cannot be talked into agreeing with the code.
Related
- MCP Server, where an independent oracle checks an AI interface's answers against the raw data
- AI Restaurant Receptionist and Keepford, the same principle in voice ordering and phone booking
- TargetMobi, the SMS survey platform for research and healthcare programmes where I was CTO from 2013 to 2018
State as of 4 October 2026: the live demo runs build c374d82, which differs from 9721cf6 only by link-preview metadata in the page head; the evaluation evidence is commit ee47d7f, measured at 9721cf6. Evaluation figures come from the project's generated evaluation report.
Facing a build like this?