The business caseThe problem this solves
A subcontractor works a day of time-and-material under a change directive and submits a ticket: crews by badge and classification, hours straight and overtime, equipment hours, materials by delivery ticket, each line priced. The field desk countersigns it — "for time only" — and it goes toward a pay application. Three other records already say what the day was worth: the general contractor's own daily field report for that date, the directive log that says whether the directive was open and whether it authorised overtime, and the contract rate schedule. They disagree with the ticket in expensive ways — hours the report does not record, a line the sub foreman struck in a note, an entry a correction re-coded to base scope, a shift already billed last week — and the countersignature printed on every ticket in this corpus is wrong on 48 of the 62. Opening one ticket packet, reading every ticket line against the day's report entry by entry and badge by badge, checking the directive was open on the work date and authorised any overtime billed, pricing each supported quantity at the schedule's own rate, reading each field note to see whether it strikes a line, re-codes an entry, records one as already billed or says another badge covered the shift — and writing down which lines are not supported before the ticket is billed.
Audience
The general contractor's field desk and project accounts, working a period's T&M tickets before they reach a pay application — and the project manager who has to answer the subcontractor when a line is named. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual ticket packets
The corpus is 62 ticket packets, 0.20 MB (txt 62). It is generated because it has to be. A real month of T&M tickets is a general contractor's commercial record: subcontractor names, crew badges that map to real workers, negotiated labour and equipment rates, and field notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that strikes a line or re-codes an entry — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 62 ticket packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every project, subcontract, ticket, directive, rate and delivery ticket is invented, a worker is a badge number and a classification, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note is from the sub foreman, the field desk, project accounts or the owner's representative. evals/check_labels.py sweeps all 62 files for an honorific and a two-word capitalised name on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your ticket packets. That is the whole change — there is no database to migrate.
==============================================================================
TIME-AND-MATERIAL TICKET VERIFICATION -- ONE TICKET, ONE WORK DATE
==============================================================================
FILE TMT-0001
PROJECT P-2207 parking structure, levels 1 to 4
SUBCONTRACT SC-31 electrical
TICKET T-5101
WORK DATE 2026-06-02
DIRECTIVE CD-05
RATE SCHEDULE RS-31-2026
TOLERANCE 10.00
CURRENCY USD
-- CHANGE DIRECTIVE LOG (this subcontract) -----------------------------------
DIRECTIVE ISSUED CLOSED OT AUTH DESCRIPTION
CD-04 2026-05-03 2026-05-30 NO earlier electrical change, level 3
CD-05 2026-05-27 -- YES electrical T&M work, level 1
CD-06 2026-06-03 -- NO added electrical scope, level 2
-- CONTRACT RATES (RS-31-2026, as attached to the subcontract) ---------------
CODE RESOURCE UNIT ST RATE OT RATE
FM foreman electrician HR 112.00 168.00
JW journeyman electrician HR 98.50 147.50
AP apprentice electrician HR 61.00 91.50
SL scissor lift 26 ft HR 18.50 --
BL boom lift 45 ft HR 42.50 --
M-EMT1 1 in EMT conduit, 10 ft EA 21.50 --
M-WIRE12 12 AWG THHN wire, 500 ft reel EA 89.00 --
-- T&M TICKET AS SUBMITTED ---------------------------------------------------
LINE TYPE RESOURCE CODE DESCRIPTION QTY UNIT RATE AMOUNT
TL-01 LAB E-125 JW journeyman electrician, ST 8.0 HR 98.50 788.00Abridged — the file continues.
The outcomeWhat a good result looks like
One ticket packet in, one row out: which lines are still claimed, which report entry supports each line with the entry's row quoted verbatim, the claimed, supported and unsupported amounts to the cent, every unsupported line named, and one of five FTV-2026 verdicts. 56 of 62 tickets come back right on all six graded fields, against 40 for the best free arm and 3 for the countersignature the field desk prints today.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, nothing stopped at the ceiling and no call failed. The 6 tickets it got wrong are named in the kit README with what it answered: 3 counted a report entry a field desk correction had re-coded to base scope, 1 left a claimed line out after a field note told it to reject that line, 1 returned no pairs on a closed directive, and 1 quoted a report row that is not in the file. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your field desk records corrections as a changed cost code, strikes lines on the ticket itself, and your billing system refuses a report entry already on a paid ticket — the free column join, and do not buy a call at all
40 of 62 tickets for $0.00. Short hours, short materials, off-rate lines, unauthorised overtime, base-coded entries, the wrong day's report, a closed directive and a report with nothing on it are all decidable from columns, dates and the registers, and the free join gets 40 of those 40. - Strikes, re-codings, already-billed shifts and covering badges arrive as free text — a field note, a foreman's message pasted into the packet — the paid call
This is the whole product. On the 22 tickets where a sentence decides the reading the paid arm is 18 and the free join 0. - A field note carries an instruction — reject a line, bill the owner, change the report — the free column join, and read the paid arm beside it
On TMT-0060 the paid arm silently left the line the note said to reject out of its reading, and an OFF-RATE ticket came back SUPPORTED. The free join reads no note and cannot be steered by one. - You want the field desk's countersignature audited — either paid or free — both beat it comprehensively
The countersignature is SUPPORTED on all 62 tickets and disagrees with FTV-2026 on 48. It is published as an arm precisely so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Your tickets span several work dates, or a subcontractor's own daily log sits beside the field desk's report — neither, yet
No packet in this corpus does. The unit of work is one ticket for one date against one report, and every percentage on this page is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own ticket packets in the same shape, data/tickets.json with your own register and data/rates.json with your own schedules, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per ticket for arithmetic you already have. That is the case against the best-fitting scenario (“Your field desk records corrections as a changed cost code, strikes lines on the ticket itself, and your billing system refuses a report entry already on a paid ticket”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A daily field report that records crews as a head count rather than by badge. The key pairs a labour line to the entry for its own badge, or to the covering badge a note names, and the free join pairs on the badge alone; a report with no badges gives neither anything to pair on. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-tm-ticket-verify. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the live button and a new scored run.






