The business caseThe problem this solves
A carrier's fuel and toll desk gets a fuel-card statement for one vehicle every fortnight and has to agree it with three records the card issuer never sees: the tank the asset register rates the unit at, the route log for each of the fourteen days, and the card-assignment register. The carrier's own fleet system already prints an exception report from those COLUMNS — a rating on a register row, a milepost on a log line, a card number on an assignment. What decides the answer is often a sentence: a saddle tank replaced before the period and not yet re-rated, a diversion the telematics never recorded, a leg run on a day the log shows no row, a card reassigned with an effective date, a fill the issuer reversed. The fleet system's own report disagrees with the procedure on 21 of the 64 statement periods in this corpus. Opening one statement period, checking every fill against the tank capacity actually in force, against the odometer gap since the last fill, against the corridor and milepost range the unit ran that day and against the card the register puts on the unit on that date, and reading every desk note for a tank replacement, a diversion, a telematics gap, a card reassignment or a reversal that moves one of those four readings.
Audience
The fuel and toll desk at a regional carrier, working a fortnight's statements before the card invoice is approved — and the route, depot and card desks an exception is handed to. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reconciliation packs
The corpus is 64 reconciliation packs, 0.32 MB (txt 64). It is generated because it has to be. A real fuel-card book is a carrier's own commercial record: card numbers, site contracts, negotiated rates, telematics traces that locate a named driver on a named road at a named minute, and desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 reconciliation packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every carrier, depot, corridor, site, vehicle and card is an invented code, every transaction, route leg, assignment and note is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a driver is a resource code and every note speaks for the fuel, route, depot or card desk. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your reconciliation packs. That is the whole change — there is no database to migrate.
================================================================================
FUEL CARD RECONCILIATION -- ONE VEHICLE, ONE STATEMENT PERIOD
================================================================================
FILE FCR-0001
VEHICLE V-3100 class TRACTOR-A
DEPOT DEPOT 2
STATEMENT ST-20260614-0001
STATEMENT PERIOD 2026-06-01 to 2026-06-14
STATEMENT CLOSE 2026-06-14
PACK COMPILED 2026-06-24
CURRENCY USD
-- VEHICLE AND TANK FACTS (fleet asset register) -------------------------------
RATED TANK CAPACITY (GAL) 150.00
RATED MILES PER GALLON 6.5
ODOMETER AT PERIOD OPEN 150,000
-- CARD ASSIGNMENT REGISTER (exported 2026-06-15) ------------------------------
CARD VEHICLE FROM TO
CD-41000 V-3100 2026-01-05 --
CD-41431 V-3411 2026-02-03 --
-- ROUTE LOG (telematics, one row per day the unit reported) -------------------
DATE CORRIDOR FROM MP TO MP MILES ODO START ODO END
2026-06-01 CR-12 20 400 380 150,000 150,380
2026-06-02 CR-12 400 870 470 150,380 150,850
2026-06-03 CR-12 870 310 560 150,850 151,410
2026-06-04 CR-12 310 690 380 151,410 151,790
2026-06-05 CR-12 690 220 470 151,790 152,260
2026-06-06 CR-12 220 780 560 152,260 152,820
2026-06-09 -- -- -- 0 152,820 152,820
2026-06-10 CR-12 780 400 380 152,820 153,200
2026-06-11 -- -- -- 0 153,200 153,200
2026-06-12 CR-12 400 900 500 153,200 153,700Abridged — the file continues.
The outcomeWhat a good result looks like
One statement period in, one row out: the tank capacity in force for every transaction, the route actually run on each of the fourteen dates, which transactions are on an assigned card, which the issuer reversed, every transaction's flags, the exception set in ladder order and one of six FCR-2026 verdicts. 60 of 64 periods come back with all seven graded fields right after the pure-code station re-derives them, against 45 for the best free code and 0 for the fleet system's own exception report.
And when it cannot
⚠︎ AND WHAT IT DOES WHEN IT CANNOT, WHICH IS THE HALF A SINGLE NUMBER HIDES. The arm's OWN arithmetic — the flags, the exception set and the verdict it wrote for itself — is right on 27 of 64. The same replies with those three fields discarded and re-derived in code from the arm's four readings are right on 60. THE MONEY BUYS THE READING, NOT THE ARITHMETIC. On the scored run 63 of 64 replies parsed; the 64th (FCR-0031) ran to the output ceiling and is counted WRONG inside the denominator, never dropped and never re-fired. The three other misses are named in the kit README with what it answered: FCR-0039 returned an empty assigned-card reading on a clean statement, FCR-0062 kept a card assignment that had lapsed, and FCR-0002 invented a leg on a date the route log has no row for.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your fleet system records tank replacements, diversions, telematics gaps and card reassignments as a STATUS on the row — the free modal floor, and do not buy a call at all
43 of 64 packs for $0.00, and on the 37 packs the columns decide it gets 37 against the paid arm's 33. Over-tank, twice-in-range and off-route are all decidable from columns and dates. - Tank replacements, diversions and card reassignments arrive as free text — a fuel desk note, an email pasted into the pack — the paid call
This is the whole product. On the 27 packs where a sentence decides a reading the paid arm is 27 of 27, the floor of record 13 and the modal floor 5. On the 21 packs where the fleet system's own report disagrees with the procedure it is 21 of 21 against 9. - You want a real exception never quietly closed as RECONCILED — the paid call, and watch kept_a_lapsed_assignment
The modal floor returns RECONCILED on a flagged statement 21 times; the paid arm does it once (FCR-0062, a lapsed card assignment kept). That is the costly direction and it is published as a count, not an average. - Your asset register is authoritative and never lags a tank replacement — the paid call with
capacity_formoved out of the readings and into the columns
That is the one reading a false sentence moves — 6 of 6 under the adversarialtankclause — and it is NOT separable from free code anyway (p = 0.375). Taking it away from the model removes the kit's whole attack surface at no measured cost.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own statement packs in the same six-block shape and data/vehicles.json with your own asset and card-assignment register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per statement for arithmetic you already have. That is the case against the best-fitting scenario (“Your fleet system records tank replacements, diversions, telematics gaps and card reassignments as a STATUS on the row”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A statement whose transactions carry no odometer reading. FC-2 is an odometer-gap test against the range the earlier fill bought; with no odometer there is nothing to test and every twice-in-range exception is lost silently. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-14 — r002-fuelcard-recon. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run, and scores the free arms offline. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run; with the key blanked the board's own /api/model route answers HTTP 400 and says so, which is the frame the empty screenshot was taken from.









