The business caseThe problem this solves
A customer asks about an order that has not completed, and the reply is not a lookup. Two things go wrong independently. The customer often is not still asking what they opened with -- 23 of these 48 threads pivot part-way through -- so an assistant that matches on the whole conversation answers the question that was abandoned. And the committed due date printed on the order is ALWAYS clean: on 24 of 48 threads it is a date nothing on the order will meet, and the only thing that says so is a milestone state or one sentence in the provisioning notes. The dangerous failure is therefore never an obviously wrong answer. It is a confident date, taken off a field that is right there in the header, given to a customer who will plan around it. Somebody opening the order, reading the thread to work out what is actually being asked, and deciding whether the date on the header is one they are willing to repeat.
Audience
Care and order-management staff who answer in-flight order enquiries, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual customer inquiry threads, each with its order record
The corpus is 48 customer inquiry threads, each with its order record, 0.07 MB (txt 48). A real order-status inquiry is a named customer, a service address, a phone number and an account number attached to a complaint. There is no public corpus of those and there should not be. Everything the measurement needs is structure -- a thread whose last question is not always its first, a printed date that is sometimes superseded, and a note that invalidates a date every field still shows as firm -- and structure can be built honestly.
The corpus
- The 48 customer inquiry threads, each with its order recordgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your customer inquiry threads, each with its order record. That is the whole change — there is no database to migrate.
Order Status Inquiry
================================================================
Thread OS-0001
Opened 2026-08-27
Inquiry Thread
----------------------------------------------------------------
[2026-08-27 12:12] customer: Can you give me a date for the install?
[2026-08-27 12:22] agent: Bear with me a moment while I look at the record.
[2026-08-27 12:35] customer: Any update? When is this actually going in?
Order Record
----------------------------------------------------------------
Order ORD-64440
Product Business Fibre 1Gb
Placed 2026-06-14
Committed due date 2026-09-19
Order status IN_PROGRESS
Appointment 2026-10-04 13:00-17:00
Jeopardy flag NONE
Milestones
----------------------------------------------------------------
2026-06-14 ORDER_ACCEPTED COMPLETE
2026-06-18 SURVEY COMPLETE
2026-06-21 BUILD COMPLETE
2026-08-19 RESCHEDULE BOOKED
2026-09-19 APPOINTMENT BOOKED
Provisioning Notes
----------------------------------------------------------------
Survey attended 18 June. Route confirmed from the footway chamber on Ferndale Way to the customer demarcation point.
Nothing outstanding on the provisioning side. Order is tracking to the committed date.
Appointment moved to 04 October 2026 at the customer's request on 19 August.
Account
----------------------------------------------------------------
Account holder Ashfold Veterinary Group
Service address 50 Ferndale Way
Contact 07700 900904
Billing account BA-661162
The outcomeWhat a good result looks like
Per thread: the operative intent, a three-way date commitment (COMMIT / WITHHOLD / HAND_OFF), the date when one may be given, the blocker when one may not, who has to act next, and the two or three sentences an agent would actually send. Beside every line, what the strongest free code would have answered.
And when it cannot
The model handed off 6 of 10 threads that needed a person, and the free floor handed off 10 of 10. On all four misses it named PORT_REJECTED correctly and then classed the thread WITHHOLD rather than HAND_OFF. That is a defect in src/prompt.py, not in the model: the rule 'a rejected port goes to a person' lives in the answer key and the prompt only gestures at it. It is recorded and NOT fixed, and r001 was not re-fired -- a prompt edited after reading the misses and then re-scored is a number fitted to its own answer key. Separately, 7 of 192 cells are an ANSWER-KEY defect rather than a model error: the key derives 'who acts next' from the blocker alone, so a cancellation hand-off with no blocker is keyed NOBODY, and the model answered CARRIER -- which is the better answer. Diagnosed, published, unfixed.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Everything that stops a date is already a STATE -- a milestone code, a jeopardy flag, a hold_reason field. — the free floor -- milestone-gate, $0.00
It takes 100 pct of structured blockers, 100 pct of hand-offs, 100 pct of dates including every superseded header, and raises no false withhold. It is free and it beats the model on hand-offs. - Field engineers, permit desks and access suppliers write what they found into a free-text note, and that note is where the reason lives. — the model, 0.00269 USD per thread on the projected card
12 of 12 prose blockers, 0 false commitments in 30, 0 over-withholds in 18 -- including the 6 threads whose note describes a problem that was FIXED, which is what an over-cautious arm trips on. - The customer changes what they are asking mid-thread. — read the last turn -- free
Floor 2 is a five-line change to floor 1 and takes the superseded-intent column from 47.83 to 100 pct on this corpus. - Somebody can write into the notes field. — the free floor, or the model with the notes gated
One instruction-shaped sentence appended to the notes turned a correct withhold into a committed date on 1 of 24 threads (x001). The free floor is immune by construction because it never reads notes.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | tools/build_corpus.py writes data/corpus/*.txt and data/gold.jsonl. None of the measured figures on this page transfer to your own queue. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it where any decisive evidence is a sentence: it catches 0 of 12 prose blockers and commits to a date on every one of them. That is the case against the best-fitting scenario (“Everything that stops a date is already a STATE -- a milestone code, a jeopardy flag, a hold_reason field.”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A file layout that is not this one. src/record.py is three regular expressions written for these headings and this two-space field separator; point it at another carrier's export and it returns empty fields -- loudly, as absent, never as a guess. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A SECOND TIER WAS NEVER RUN. One model answered every scored thread. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-27 — r001-order-status-ask. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Every thread, the answer key, all three free floors, the injection probe's result, the token split and the recorded run all ship in the repo. python3 -m src.app renders the whole page with no key and no network.


