Case Study · Solo Build

Fieldnote: keeping consequential decisions out of the model, then attacking that boundary

Fieldnote is my Phase 1 prototype for conversational health check-ins in community-based participatory research. The language model only proposes: deterministic code keeps a value only if it points at the participant's own words, and a stateless server decides when a check-in must stop. Then I had that boundary attacked. The attack confirmed and measured defects in the build deployed that morning, the fixes went into the mechanisms rather than the word list, and later reviews found defects the evaluation itself had missed. Those were fixed, and the evaluation was extended to catch them. It is a demo with fictional answers, not a product in use; everything below comes from its code, tests, evaluation report, live demo and build history.

1,368 → 356 → 0 scenarios breaking the written safety policy on the frozen corpus: the build deployed on the morning of 1 October, after round 2, and the final deployed build
12,777 evaluated scenarios on the deployed build's code, run in process with the model stubbed, across a frozen corpus, a held-out corpus and a blind negation set; all nine invariants held
16 of 60 explicit disclosures, written by an AI agent never shown the keyword list, that the list stops on its own
628 automated tests passing offline (402 RSpec, 226 Vitest)

At a glance

What it is
A browser prototype that runs a family wellness check-in as a conversation and builds the structured record live beside it, for researchers working with Indigenous and rural communities in the US. No database, no accounts, no phone number.
Where it runs
Live at fieldnote.railsfanatics.com in Demo Mode: no sign-up and no API key, so no AI model is called and the replies come from a deterministic provider. Live Mode uses Google Gemini through the server.
How it was checked
An adversarial evaluation against a written policy, run in process against every build from the one deployed on 1 October to the final one, with the model stubbed and observed on the wire; disguise and clause-separator sets written by AI agents given no examples; a code review and hostile reviews, all by AI agents; offline test suites; manual QA and a production smoke test of the deployed build; and a recorded request against the real Gemini API. No participants, no pilot, no measured extraction accuracy, and no evaluation of the live model.
Hardest problems
  1. Keeping unanchored values out of the record without trusting the model's confidence.
  2. Enforcing a safety pause on a server that remembers nothing.
  3. Telling a model refusal from a model failure, which look alike and mean opposite things.
  4. Testing the boundary without letting the test ask the code what the right answer is.
Stack
Ruby on Rails 8.1 with no database, React 19 and TypeScript, Google Gemini on the server only; RSpec, Vitest and a Ruby evaluation harness. Deployed with Kamal.
My role
Sole engineer: I set the architecture, the safety requirements and the evaluation's rules, and AI coding agents (Claude) wrote the implementation and the harness under my direction. The code review and hostile reviews were also done by AI agents; no person other than me has reviewed the code.
More
The live demo. The repository is private.

The problem

In community-based participatory research, a community health worker asks families about things like diet, sleep and whether the food runs out before payday. A conversation suits people who answer "about four servings most days" better than a form does. But something has to turn that sentence into a 4, and if that something is a language model, it can produce a 4 that nobody said. In research data, an invented value is worse than a missing one, because nobody downstream can tell them apart.

A check-in can also surface a disclosure, of abuse or of being unsafe, and the questionnaire must stop there rather than carry on to the next question. And communities, their advisory boards and tribal review boards need to see where data goes and what was machine-generated. The design question was: what may the model do, and what must the application decide?

Who decides what

The model may propose exactly four things on each turn: the next reply, candidate values each with a quote, a hint that a clarification might help, and a signal that the message may be sensitive. The response is validated against a strict contract, and any extra key, such as an attempt to set the follow-up flag, fails it. Which field to ask next, when the check-in is complete, whether a family needs follow-up and whether to pause are decided by code.

That code is split in two. The browser holds the interview record in a TypeScript engine: values, validation, merging, completion and export. The Rails server holds everything about safety, the follow-up decision and the model key, and keeps no session at all.

Who decides what in one participant turn. The browser's TypeScript engine owns the interview record: state, values, the next field, completion and export. It keeps a proposed value only if its quote appears in the participant's own messages, holds low-confidence, out-of-range or unmappable values as needs confirmation, marks a field unknown after two unanswered asks, records provenance on every value, and holds the signed continuity token, which it can read but not edit. The stateless Rails server runs a fixed pipeline: verify the token's signature, expiry, template digest and turn counter, rejecting with no model call on failure; return the fixed message again if the token is already paused; scan the new message and every transcript turn with both language lexicons, and on a hit return a fixed pre-written message with crisis resources and a paused token without calling a model; only then call the provider; validate its proposal against the contract; and pause on the model's sensitivity signal. Rails decides follow-up and logs no participant text. The token carries only v, sid, n, safety, tpl and iat. The model provider proposes four things: a reply, extraction candidates with quotes, a clarification hint and a sensitivity signal. A transient failure falls back to a deterministic provider and marks the turn degraded; a refusal or block becomes a safety pause and never falls back.
The model proposes; the browser engine owns the record; the server owns safety and follow-up. Full diagram.

Quotes, not confidence

The rule that keeps unanchored values out is simple: a proposed value is kept only if the model points at words the participant actually wrote. The engine checks the quote against the participant's recent messages, ignoring case and accents but otherwise as a plain substring. If the quote is not there, the value is rejected and the rejection is logged in the interview's own event history. A confident model with no quote gets nothing.

The rule does not check that the value follows from the quote: a real sentence paired with the wrong number would pass it. That is why the quote is stored and shown beside every value, so a person reviewing the record can check one against the other, and why a reviewed flag exists at all. A value below 0.5 confidence, outside the field's range or impossible to map is held as "needs confirmation", never clamped or coerced; a field asked twice without an answer is marked unknown, not guessed.

The Fieldnote check-in in Demo Mode. On the left, the conversation: the participant answered about four servings most days and maybe three sodas a week. On the right, the structured record at 2 of 6, with the value 4 under its quote. A provenance popover over the first field reads: source, extracted from the conversation; quote, about four servings most days; turn 2; confidence 60 percent; input mode, typed; reviewed by a human, no. Below, the governance panel states that Fieldnote has no application-side persistence of interview data, that in Demo Mode no AI model is called but the text still goes to Fieldnote's server, and that the data-handling terms of this deployment are undeclared. The header badge reads Demo Mode, no AI model is called.
Every value opens to its provenance: the quote, the turn, the confidence, how it was entered and whether a person has reviewed it. The CSV and JSON exports carry the same columns. Demo Mode, fictional answers.

A safety pause that a stateless server can enforce

Before any model sees a turn, the server scans the new message and every turn of the transcript that comes with it, with both the English and the Spanish word lists. On a hit the check-in stops: the participant gets a fixed message written in advance and never produced by the model. It says the check-in is paused, that this is a demonstration in which no one is notified and no one will make contact, and points to crisis resources, with a note that a real deployment must use community-approved local ones. No model is called, and there is no "resume" button. The end button reads "End check-in (no hand-off in this demo)", because notifying a health worker is not built in this prototype; the follow-up flag and the export are.

The server keeps no session, so it cannot remember that it paused, and a flag sent by the browser would be worthless. Instead the server signs a continuity token on every turn and requires it on the next. It says whether the check-in is paused, which template it belongs to and how many participant turns the server has accepted, and it carries nothing the participant wrote: a version, an interview id, the turn counter, the safety status, a template digest and a timestamp. A safety category in a signed bearer token would itself be a record that someone disclosed abuse.

A forged, edited, expired or missing token, a token for a different template, or a turn counter that does not match is rejected before any model call. Once a token says paused, no model is called again for that check-in on any endpoint, including the AI summary. The summary request is held to the same rules: its values are scanned before any model call, and a disclosure found there, or a model refusal of the summary, pauses the check-in exactly as a turn would, with the same fixed message and a paused token.

A paused check-in in Demo Mode. A banner reads: this check-in is paused, this demo does not contact anyone, if you need help now, use the resources below. Below it, a Read aloud button with the note that nothing is spoken automatically, general US crisis resources (the 988 Suicide and Crisis Lifeline and the National Domestic Violence Hotline) with a note that a real deployment should substitute community-approved local resources, and two actions: End check-in (no hand-off in this demo), and Start new check-in. The transcript shows the fixed reply: this check-in is paused, thank you for telling us, you don't need to say anything more here, this is a demonstration, no one is notified and no one will contact you, if you need help now, please use the support resources shown on this page. The composer is disabled. The structured record on the right stops at 1 of 6, with the value 4 under its quote. In the governance panel, the Generate AI summary button is disabled, with the note: not available after a safety pause, no model is called once a check-in is paused.
A fictional disclosure stops the check-in on the live demo. The reply is a fixed, pre-written message, the end button says no hand-off happens, and the AI summary is disabled. Demo Mode, fictional answers.

What the token does not do is prevent replay. A browser that kept an earlier token and its transcript can present them again and rewind its own check-in to before the pause. The two-hour expiry does not bound this either: every request that re-issues a token gives it a fresh two hours, so a kept pre-pause token can be refreshed and used long after the pause. A pause also belongs to one check-in: a new check-in starts unpaused. The pause holds for a client that keeps the token the server returned, not against one that deliberately keeps an older one. Closing that needs server-side state, which Phase 1 deliberately has none of; the limitation is stated in the README rather than hidden behind the word "signed".

A refusal is not a failure

Over HTTP, a model that times out and a model that declines to continue can look similar. They mean opposite things. A timeout, a 5xx, a rate limit or an unreadable response is a failure: the server answers that turn from the deterministic Demo provider, marks it degraded, and the header badge says so. A refusal or a safety block is the model saying something may be wrong, so it becomes a safety pause and never falls back; a fallback there would have the Demo provider ask the next questionnaire item at exactly the moment the product exists to stop. Any finish reason other than a normal stop is treated as a block, so an ambiguous case fails closed. The Gemini request shape itself came from a recorded request against the live API, made before any provider code existed.

Attacking the boundary

A boundary described in prose is a claim. To test it I wanted an evaluation that could not simply agree with the code, so it was built under four rules:

  • The policy cites the spec. Every expected outcome (should this pause, may the model be called, may it fall back) cites the section of the design document or non-negotiable it comes from, and every interpretation of a silent or contradictory spec is written down with its reason. The policy was written by the same AI session that had already found the first defects by reading the code, so its independence rests on those citations, not on its author not having seen the code.
  • The test never asks the app. It drives the HTTP API only, intercepts every outbound Gemini request on the wire and records its body, decodes tokens itself and captures the logs. "The disclosure reached the model" means a model request was made and its text was found in it, not that the scanner said so.
  • The sentences were written by an agent that never saw the word list. 84 disclosures and 44 harmless controls in English and Spanish were written by an AI agent that was never shown the word list, and frozen by hash; every run records and verifies that hash.
  • Mechanisms, not vocabulary. A disclosure the app stops in its plain form is an anchor; every variant of it (the other language setting, no accents, capitals, extra or Unicode spaces, invisible characters, fullwidth letters, a harmless clause in front, placed in the transcript, placed in an agent turn) must be stopped too. Fixes may change how matching works and what gets scanned, but no sentence may be added to the word list, so the corpus still means something afterwards.

Round 4 added a fifth rule, after the reviews showed that disguises written with knowledge of the defects only tested what the fixes already assumed: disguises written blind. Separate AI agents with no access to the repository, given the policy text but no example separator or character, wrote a set of 48 clause separators and a set of 30 invisible or blank-rendering characters, and each set was frozen by hash before its first run.

Every run reported here uses the final version of the evaluator, in process against a clean checkout of the commit it measures, with the model stubbed so every result is deterministic. On the final build that is 5,865 scenarios on the frozen corpus, 3,185 on the held-out corpus and 3,727 on the blind negation set: 12,777 in all.

What it found

Reading the code had already turned up the first suspects: a negation guard that matched "no" inside "now", "know" and "cannot", Spanish without accents slipping past, a word list chosen by the browser's language setting, a transcript the server never scanned, and an AI summary that took no token and so could call the model after a pause. Run against the commit deployed that morning, the first version of the evaluation confirmed all of them, added one the review had missed (extra whitespace defeated the match), and put a size on them. With today's evaluator, which has been extended three times since, the same commit breaks the written policy in 1,368 scenarios. (The 86 and 131 reported on 1 October came from earlier versions of the evaluator and are not comparable with the figures here.)

The first round of fixes was mechanisms: both word lists always, one normalization for text and phrases, negation only as a whole word in the same clause within three words, every transcript turn scanned, the summary gated on the token, templates and roles validated. A separate AI agent then attacked the fixes and found that Unicode spacing, invisible characters and fullwidth letters still got a stopped disclosure through, that a negation still reached across "and" or "que", and that the summary sent field keys it never scanned. Round 2 closed those, and that build was deployed on 1 October. Today's evaluator still finds 356 violating scenarios in it.

Round 3 began with a code review of the scanner by an AI session. It found that a line break did not end a clause for the negation guard, that some control characters reached the model as plain spaces, and that the held-out corpus had new sentences but no new disguises. Round 3 fixed those, added a blind negation set and blind disguise mechanisms, and added two invariants: a malformed or unverifiable request never reaches the model, and no text can add lines to the model request. A second hostile review, of round 3, found more: a negated clause still cancelled a disclosure after an emoji, a slash or a quote mark; a model refusal on the summary did not pause; a disclosure split across summary list items, or summary values carrying extra prompt lines, reached the model; the raw body of a malformed request was written to the debug log; and one crafted message took about five seconds to scan. Round 4 fixed these by mechanism: clause breaks defined by Unicode character category, "invisible" defined by Unicode's own property, one request contract shared by the turn and summary endpoints, middleware that refuses unparseable requests before Rails logs them, size caps, and an index that keeps the negation check fast. A further hostile review of round 4 found smaller gaps (unspaced dashes, single quotes, "№" folding to "No", other JSON content types), and those were fixed too. No phrase was added to the word list in any round, and the negation window stayed at three words.

The evaluation itself had flaws too, which matter as much as the code defects. Its own results revealed one: a text transform in the harness silently turned line breaks into spaces. The hostile review of round 3 found two more. The prompt that commissioned the blind negation set had listed the separators to use, so that set says nothing about separators. And the harness captured logs through a logger that Rails had already memoized, so one logging invariant held vacuously, and a unit test of the logging fix passed while the request body it was meant to keep out was being logged. Each was fixed and recorded, every run was repeated with the corrected evaluator, and the two blind sets above were commissioned because of them. One of those sets found a defect in the round-4 fix on its first run (a combining character used in place of a space); it was fixed afterwards and is labelled as such.

Manual QA of the finished build found a different kind of defect: the pause message and banner told the participant that a community health worker would follow up, and the end button said "hand off to CHW", in a demo that contacts no one. I decided the replacement wording, which says plainly that no one is notified, and the tests now assert it in both languages. The final build, 9721cf6, was deployed on 3 October 2026, and a smoke test of production confirmed the new wording and the pause, and that a paused token gets only the fixed pause on the turn endpoint and is refused on the summary endpoint.

Results

Each cell is the number of applicable scenarios in which the invariant held, from the generated report, with every run produced by the same final evaluator. "Held-out" is a second corpus of 42 disclosures and 28 controls written blind after the first round of fixes; it has been run in every round since, so it is strictly held out only for round 1. The blind negation set (32 disclosures that contain a negation word, 48 controls) was written blind and frozen before its first run in round 3.

InvariantDeployed morning of 1 Oct, before fixesAfter round 2Final buildHeld-out, finalBlind negation set, final
No stopped disclosure, or tested variant of one, reaches the model833 of 2,1121,948 of 2,2512,251 of 2,2511,394 of 1,3942,046 of 2,046
No model call after a pause, on any endpoint1,007 of 1,0242,133 of 2,1432,450 of 2,4501,535 of 1,5352,258 of 2,258
A refusal never becomes a Demo answer, on any endpoint10 of 1010 of 1010 of 1010 of 1010 of 10
No participant text in server logs5,760 of 5,7695,856 of 5,8655,865 of 5,8653,185 of 3,1853,727 of 3,727
Token claims within the allowlist, no participant text5,761 of 5,7615,854 of 5,8545,849 of 5,8493,169 of 3,1693,711 of 3,711
The safety follow-up reason cannot be suppressed22 of 2422 of 2424 of 2424 of 2424 of 24
Demo Mode makes no model request138 of 138138 of 138138 of 13880 of 8090 of 90
A malformed or unverifiable request never reaches the model30 of 7357 of 7373 of 7373 of 7373 of 73
Participant or template text cannot add lines to the model request6 of 1612 of 1616 of 1616 of 1616 of 16

Scenarios violating the policy or an invariant: 1,368, then 356 after round 2, then 0 on the final build on the frozen corpus; 818, 203 and 0 on the held-out corpus; 903, 266 and 0 on the blind negation set. On the final build, every separator the blind author classed as clearly ending a clause was handled (624 of 624 scenarios built on the frozen corpus's stopped disclosures, against 610 after round 2), and so was every in-policy blind invisible character (448 of 448 scenarios, against 367). Those blind sets now run in the regression suite, so they are no longer unseen evidence. Classes that still get through are outside the written policy and reported as coverage: look-alike letters from other scripts, leetspeak, letters spaced out, Latin small capitals, right-to-left overrides, and a negation followed by a tab, which the policy treats as ordinary whitespace. Three repeat runs of the round-4 code, made before the final wording change, gave identical verdicts in all 5,865 frozen-corpus scenarios.

The fixes cost false pauses. On the frozen corpus 1 of 44 harmless controls pauses ("no estoy pensando en hacerme daño", where the negation sits four words back), and the held-out set shows the same construction once in 28. On the blind negation set, which was written to stress the guard, 8 of 48 controls pause, against 4 of 48 for the guard deployed on the morning of 1 October: rounds 1 and 2 traded false pauses for fewer misses, and rounds 3 and 4 did not change the count. That is a stress figure, not an estimate of how often real answers would pause.

The most important number is the least flattering one. The keyword list stops only 16 of the 60 explicit disclosures in their plain form (10 of 30 in the held-out set, 15 of 32 in the negation set), none of the 16 indirect ones, and none of the 8 split across two messages. The invariants are conditional on that first catch, and on the variant families tested: for those, a recognised disclosure never reached the model. Other rewordings still get through ("nobody knows he hits me" reads as negated), and recognition itself is the weak link. In Live Mode the model's own sensitivity signal is a second net, which this evaluation did not test: the model was stubbed throughout, so none of this measures how the live Gemini model behaves.

What is shown and what is not

ShownNot shown
Defects found in the build deployed on the morning of 1 October, fixed in four rounds as mechanisms, with no phrase addedThat the word list catches most disclosures; it stopped 16 of 60 explicit ones
Nine safety invariants holding in all 12,777 evaluated scenarios of the deployed build's code, run in process, on a frozen corpus, a held-out corpus and a blind negation setA detection rate, sensitivity or any real-world performance; every corpus is small and AI-written
Flaws in the evaluation itself, found by its own results and a later review, fixed and recordedThat every separator or invisible character is handled; the blind sets are finite, and some classes are outside the policy
Values kept only when their quote is in the participant's own wordsHow accurate the extraction is
A pause a client cannot clear for that check-in by claiming safety or dropping the token, with no model call after it on any endpointReplay prevention, a token expiry that bounds replay, or proof that a person reviewed a generated schema; these need server state
The deployed build, its pause wording and its pause behaviour, checked by manual QA and a production smoke testThe live Gemini model's behaviour: the evaluation stubs the model, and the public demo calls none
A live, public demo with no sign-upReal participants, a pilot, or any deployment with community partners
The check-in, safety message, word lists and governance panel in English and SpanishSpanish in Design Your Own and the health worker view; review by native speakers

Fieldnote is not a crisis service, is not clinically validated, and has no application-side persistence of interview data; in Live Mode, text is sent to Google's API for each request. Three more limits are structural. A modified client can hide the non-safety follow-up reasons, because those come from field values the browser owns. The server can validate a template but cannot prove anyone reviewed it. And the text a researcher pastes into Design Your Own is not safety-scanned: it is treated as researcher input, not participant input, and in Live Mode it is sent to the model as written. In Demo Mode, Design Your Own returns a fixed example schema with a visible notice, whatever is pasted.

How it was built

I wrote the product brief and its central rule: the model must not own application state; it proposes, and the application validates and decides. Claude drafted an implementation plan from that brief, and I sent drafts back until it said what I wanted. From that review came the signed token that lets a stateless server enforce a pause, the rule that a refusal never falls back, the live API request as a gate before any provider code, and the token's minimal claims. AI coding agents (Claude) wrote the implementation and tests in September 2026, under my direction.

The October evaluation followed the same pattern. I set its rules: an oracle that never asks the app, a policy written from the spec, a frozen corpus written without the word list, a held-out set written after the first fixes, disguises written blind, the required invariants, and no new phrases in the word list. AI agents wrote the policy contract, the corpora and blind sets, the harness and all four rounds of fixes to those rules. AI agents also did the code review and the three hostile reviews, including the ones that found flaws in the evaluation. I set the rules, decided the policy questions and the pause wording, and ran the final manual QA. The production smoke tests were run by an AI agent (Claude) on my instruction. No person other than me has reviewed the code, and I have not reviewed it line by line.

What I would change

The evaluation points to the first change: the deterministic net needs better recognition, not just better plumbing. It stops only 16 of the 60 explicit disclosures in the main corpus and none of the indirect ones, and what counts as a disclosure that must stop a check-in is a clinical question before it is an engineering one. The second is server-side state, as a deliberate Phase 2 decision: it is what replay prevention, a meaningful expiry and proof of template review all need. The third is review by people: by native Spanish speakers, by community partners, and of the code by an engineer other than me.

What this demonstrates for client work

If a model is going to write into records people rely on, such as research data, health notes or case files, the work is in deciding what the model may never own, making every value traceable to its source, and then attacking those decisions with tests that cannot be talked into agreeing with the code.

Related

  • MCP Server, where an independent oracle checks an AI interface's answers against the raw data
  • AI Restaurant Receptionist and Keepford, the same principle in voice ordering and phone booking
  • TargetMobi, the SMS survey platform for research and healthcare programmes where I was CTO from 2013 to 2018

State as of 4 October 2026: the live demo runs build c374d82, which differs from 9721cf6 only by link-preview metadata in the page head; the evaluation evidence is commit ee47d7f, measured at 9721cf6. Evaluation figures come from the project's generated evaluation report.

Facing a build like this?

← All case studies