The business caseThe problem this solves
A support ticket arrives -- an email, a merchant-portal form, a transcribed phone call, a web chat -- and somebody on the desk decides where it goes before anybody works it: which queue this issue takes on the terminal this merchant actually has, whether this is the third call on the same box inside a month, what clock their plan buys, and whether the symptom they describe is even possible on the device on record. Most of that is a table lookup and a date subtraction. The reading in front of it is three sentences of informal prose in which "the terminal is broken" means a settlement problem, and "probably nothing" means a skimmer. The desk's cold read of each support ticket against the routing rulebook -- the queue table, the device capability check, the repeat-failure count and the plan clock -- before a person confirms where it goes.
Audience
The support desk that confirms the routing, and the operations lead who answers for a van that was booked unnecessarily or a compromised reader that was not taken out of service. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual support tickets
The corpus is 60 support tickets, 0.02 MB (txt 60). A defect mix a payments desk would recognise and no public dataset can provide: a tamper written as a spare part somebody fitted, a paper jam on a device with no printer, a battery fault on a corded terminal, a third connectivity call on the same serial inside the window beside a second one that must not escalate and a third one whose priors are ten weeks stale, an unattended kiosk with nobody on site, a merchant claiming a plan their record does not carry, a settlement problem described as "the terminal isn't working", and two tickets that say nothing at all. 16 of the 60 are deliberately clean, because a corpus that is all traps measures a different job from the one a desk has.
The corpus
- The 60 support ticketsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your support tickets. That is the whole change — there is no database to migrate.
From: Dee Storrow <storrow@example.invalid>
Subject: Terminal fault
Received: Tuesday, 08 September 2026
The handheld we take to the tables will not print a receipt -- it feeds the paper through blank. We are on your managed plan so this should be a one-hour response.
Dee Storrow
Storrow Hire Centre
The outcomeWhat a good result looks like
A routed ticket a person confirms instead of a ticket a person triages cold: the queue, the priority band, the response hours the merchant's plan buys, a verify flag when the ticket and the record disagree about the box, and one sentence naming the rule that decided it.
And when it cannot
An under-route. A tamper report sent to a field engineer leaves a possibly-compromised reader taking cards while a van is booked for next week. The free floor ships 4 of those in its 23 opportunities -- including 2 of the 5 tamper reports; the model and the recheck ship 0. What the model ships instead is 3 unnecessary dispatches in 35, which costs a day of somebody's time and not a merchant's card data.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your tickets say what they mean, your queues map cleanly from a category word, and an over-dispatch costs you more than an under-route — the free floor (evals/baseline.py) in front of the routing engine
76.7 pct of tickets all-correct for $0.00, no key, no network, and 0 unnecessary dispatches in 35. The engine, not the model, already applies every rule -- the queue table, the device override, the repeat count, the plan clock. - Tickets written the way merchants actually write them -- a skimmer described as a spare part, a settlement problem called a broken terminal, an entitlement asserted in prose — the model WITH the recheck -- the pair this kit ships
It reads 90.0 pct of issue categories against the floor's 78.3, misses no tamper, under-routes nothing, and the recheck removes the incoherent records (a refusal with a response clock attached) for no second call. It costs $0.0100 a ticket on the projection card. - You want to know whether any of this transfers to your desk — your own tickets, through the same harness, before believing any number here
The rulebook, the fleet, the bands and the history are data files; the harness, the floor and the scorer run keyless. Sixty of your own tickets with a hand-checked key would tell you more than this page can.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-30. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own tickets as .txt files into data/corpus/ and add one row per ticket to data/tickets.json -- the merchant, the plan, the installed device family and serial, the received date and the channel. ⚠︎ DO NOT point this at real tickets casually. Corpus lens → |
| When is this the wrong choice? | Avoid: Using it anywhere a tamper report can arrive. It missed 2 of the 5 on this corpus and sent both to a field engineer, and its 8 refusals are 8 tickets a person reads by hand. That is the case against the best-fitting scenario (“Your tickets say what they mean, your queues map cleanly from a category word, and an over-dispatch costs you more than an under-route”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL TICKET. The corpus is templated from small per-case phrase pools, and the floor's 76.7 pct is partly a measure of that: its keyword lists were written from the rulebook, but its phrases for "we cannot take cards" match the styles the generator writes, which is most of its win on the priority band. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether any of these rates resemble a real desk's queue. The corpus is invented, templated from small per-case phrase pools, and the defect mix is chosen; see data/SOURCES.md. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-30 — r001-terminal-ticket. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on a copy of the kit with no .env and no API key, on this machine: python3 -m evals.check_labels re-graded the key with independent arithmetic (KEY CLEAN, 60 tickets, 60 records, 26 cases, 21 planted prior tickets) and python3 -m evals.run --floor rules scored the free floor, together in well under a second; python3 -m src.app served the board with the model button disabled and saying why, and tools/shoot_ui.mjs took every screenshot with API_KEY blanked. The corpus, the estate, the history, the rulebook, the key and both recorded model runs ship in the repo, so the whole product renders before anyone decides to spend.



