The business caseThe problem this solves
A service record already says how to communicate with the person it is about — in their own words, in a note somebody took, in a letter that came back, or in the register row recording what was arranged last time. Nobody reads all five of those places for all five accommodations on every booking, so the arrangement that was made once is quietly not made again, and the request written in a sentence a checklist does not carry is not made at all. Neither failure produces a complaint that reaches anybody: what surfaces is somebody who did not attend, or who attended and got less out of it than everybody assumed. The clerk's own pass over one record's five sections — the record facts, what the person told us, the intake and contact notes, the correspondence history and the service history — against five accommodations, one at a time. It replaces the READING of that pass and none of the deciding: AP-2026 is applied in code either way.
Audience
The booking clerk or co-ordinator who decides what to book, what to send, how long to hold the slot and who else may be in the room — and who has one record in front of them and a list to get through. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual service records
The corpus is 60 service records, 0.05 MB (txt 60). The cases are the ones that make this decision hard, and none of them can be taken from anywhere real. A person who declined an interpreter last time and has since asked for one (declined_then_stated, 3 records) is the case that proves the rule ORDER rather than any rule: AP-1 over AP-3 is what makes somebody able to change their mind, and both arms get all 15 of its judgements right because the order is applied in code on both. A standing notice about the interpreter booking line, enclosed with every letter that month (phrase_false_positive, 6 records, 30 judgements), is a fact about a building and flagging it sends a clerk to arrange something nobody asked for. An interpreter asked for with the language never written down (language_unstated, 4 records) makes NOT STATED the right answer for the language field, because an interpreter booked in a guessed language is worse than an unfilled booking — it looks handled. And the whole middle of the corpus is the UNCLEAR-ASK cases — a letter that came back undelivered, a relative who did the talking, a call that moved channel — where the record points at a need and does not establish it.
The corpus
- The 60 service recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your service records. That is the whole change — there is no database to migrate.
SERVICE RECORD AR-0001 -- Booking confirmed by text
RECORD FACTS
Record opened 2026-03-17
Service outpatient booking at Tarnsyde
Contact channel used telephone
Register read as at 2026-08-31
WHAT THE PERSON HAS TOLD US
They asked whether an interpreter in Arabic could be booked for the same time slot.
They asked whether they could have a double slot.
INTAKE AND CONTACT NOTES
They said the date offered was fine and did not want to change it. They asked whether there is parking at Tarnsyde and were told where to go.
The record was opened when the referral arrived and has been added to twice since. They asked what time they should arrive and where to wait.
CORRESPONDENCE HISTORY
2026-03-31 text reminder sent -- confirmed by reply
SERVICE HISTORY -- what was arranged at earlier contacts
No accommodation has been arranged at an earlier contact.
The outcomeWhat a good result looks like
Five determinations a person confirms instead of a record they re-read: for each accommodation, INDICATED, NOT-INDICATED or UNCLEAR-ASK, with the sentence it came from copied verbatim out of the record and located at its offsets. Where the answer is UNCLEAR-ASK what it produces is a QUESTION TO PUT TO SOMEBODY with the text that raised it, so the conversation has a starting point.
And when it cannot
Three directions and they are not comparable. A MISSED ACCOMMODATION is somebody who asked, in their own words, and arrives to an appointment they cannot take part in — the free floor does this 10 times of 38 (26.3 pct), the scored call once. AN ACCOMMODATION BOOKED ON NOTHING is a wasted slot or a wasted print run on four of the five, and on accompanying_person it is a person in the room the person themselves did not choose, which is not a clerical error — floor 12, call 3. AN UNNECESSARY ASK is a person rung to answer a question the record had already settled — floor 6, call 0. And the one the kit is actually built to count: ASSUMED INSTEAD OF ASKED, a decision taken about somebody rather than put to them — floor 10 of 27 (37.0 pct), call 6 of 27 (22.2 pct). The call roughly halves it and comes nowhere near removing it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Records written in the phrases a standing checklist already carries — plain requests, plain silences, and a register that is kept — the free rules floor alone
Six results are a TIE on this corpus: stated_plain 25/25 both, nothing_indicated 30/30 both, declined_then_stated 15/15 both, third_party_remark 22/25 both, ambiguous_channel 14/15 both, and the whole clean control group 11/11 both. Every register-derived answer is free: AP-3 6/6 and AP-4 9/9 on both arms, because they come from data/records.json and no reading is involved. Zero fabricated citations and zero invented languages on both arms. - A backlog where the request is in the person's own words rather than the checklist's — the paid call, read per class
stated_in_own_words is 6 records and 30 judgements: the model takes 29 of 30 where the floor's phrase lists were written against the listed phrasings. Across the corpus the model misses 1 of 38 stated requests where the floor misses 10, and over-flags 3 of 235 silent judgements where the floor flags 18 — the two directions that cost a person their appointment or a clerk a wasted booking. hard records go 11/49 to 39/49.
And where nothing here is good enough:
- Deciding whether the call is worth its money on THIS corpus — neither arm as published — write the stem floor first
A regex reader built from the kit's own generator tables reaches 300/300 verdicts and 60/60 records for $0.00 (see baseline_note). Against that arm the paid call LOSES: 290 against 300 judgements, 50 against 60 records, 21 against 27 on UNCLEAR-ASK. The published margin measures the shipped floor's weakness and the corpus's regularity, not the difficulty of the task.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own records as .txt into data/corpus/, add one row per record to data/records.json for the register, and write one gold row per record in data/gold.jsonl with a citation_span for every judgement whose verdict is not NOT-INDICATED. ⚠︎ THE WHOLE RECORD GOES TO A MODEL PROVIDER VERBATIM — the notes, the correspondence outcomes, the register, everything a colleague wrote about a contact. Corpus lens → |
| When is this the wrong choice? | Avoid: The paid call — on this slice it buys nothing, and on AP-6 it is worse than the floor. That is the case against the best-fitting scenario (“Records written in the phrases a standing checklist already carries — plain requests, plain silences, and a register that is kept”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL RECORD. Sixty records from a generator with 82 sentence templates across five accommodations and eight mechanisms, with substituted languages and places. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | ANY SECOND MODEL, TIER OR PROVIDER. One model id, one day, one tariff window. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-access-need. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — ⚠︎ PARTIALLY EVIDENCED, AND THE REST IS NOT MEASURED. What is on disk and checkable: the 60 records, the register, AP-2026 in both renderings, the answer key with its character offsets, the free floor's recorded result, the paid run's recorded result and its full reply cache. Standard library only — no requirements to install. The free half (build_corpus, check_labels, the b000 floor, the --stub pass, the board on 127.0.0.1:9235) needs no key and no network. ⚑ THE CORPUS REBUILD WAS VERIFIED REPRODUCIBLE ON 2026-08-31 rather than claimed: tools/build_corpus.py was re-run from a clean copy under four PYTHONHASHSEED values (1, 97531, 0, 424242) and the corpus, gold.jsonl, records.json, policy.json and corpus-stats.json came back BYTE-IDENTICAL to the committed bytes every time. What is NOT measured is a cold clone on another machine, another Python, or another operating system.



