The business caseThe problem this solves
A programme's cost actuals for a month — labour from timecards, material from receipts, subcontract invoices — have to land on the control accounts their charge numbers map to before anybody can report earned value against them. The cost tool's load check reads COLUMNS: a STATUS printed on a map extract, a CLOSED on a scheduling export, the adjustments still on its list. What decides the answer is often a sentence: a map change the programme analyst withdrew after the extract, one the change board approved before AS AT that the extract still prints PENDING, a work package the scheduling desk reopened for a late receipt, an estimate reversed when the real invoice posted. The load check's own status line disagrees with the procedure on 36 of the 64 packs in this corpus. Opening one reconciliation pack, finding the map row in force for every cost actual's charge number on its posting date, reading each note to see whether it approves or withdraws a map change, reopens or closes a work package, or reverses an estimate — and when — then testing every actual against its package's status and its load, and tying every control account on the cent.
Audience
The programme-control desk at an aerospace manufacturer working a period's actuals before the control account managers review their variances — and the programme analyst, scheduling and finance desks who own the corrections an exception raises. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reconciliation packs
The corpus is 64 reconciliation packs, 0.37 MB (txt 64). It is generated because it has to be. A real programme's cost ledger is a manufacturer's own commercial record — and on a defence programme often a controlled one: named buyers and engineers on every timecard, supplier invoices, contract line items and desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies. No government earned-value management standard is cited or relied on.
The corpus
- The 64 reconciliation packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every programme is an invented code and name, every charge number, work package, control account, cost actual and estimate is arithmetic on the file index, the close calendar printed on each pack is declared fictional, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for the programme analyst desk, the scheduling desk, the finance desk, the control account desk or the programme office. evals/check_labels.py sweeps all 64 files for a person-shaped name on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your reconciliation packs. That is the whole change — there is no database to migrate.
================================================================================
ACTUALS-TO-CONTROL-ACCOUNT RECONCILIATION -- ONE PROGRAMME, ONE PERIOD
================================================================================
FILE CAR-0001
PROGRAMME PG-310 Orrindale trainer cockpit display refresh
PERIOD 2026-04 (2026-04-01 to 2026-04-30)
ACTUALS LOADED 2026-05-03
AS AT 2026-05-05
PACK COMPILED 2026-05-08
CLOSE CALENDAR corrections due 2026-05-07, acceptance due 2026-05-09 (declared for this fictional programme)
CURRENCY USD
-- CHARGE NUMBER MAP (the baseline map and every map change printed on it) -----
ROW CHARGE NO WORK PACKAGE CONTROL AC EFFECTIVE SOURCE STATUS
MAP-310-01 CN-5310-01 WP-310.1.1 CA-310.1 2026-01-05 baseline map APPROVED
MAP-310-02 CN-5310-02 WP-310.1.2 CA-310.1 2026-01-05 baseline map APPROVED
MAP-310-03 CN-5310-03 WP-310.1.3 CA-310.1 2026-01-05 baseline map APPROVED
MAP-310-04 CN-5310-04 WP-310.1.4 CA-310.1 2026-01-05 baseline map APPROVED
MAP-310-05 CN-5310-05 WP-310.2.1 CA-310.2 2026-01-05 baseline map APPROVED
MAP-310-06 CN-5310-06 WP-310.2.2 CA-310.2 2026-01-05 baseline map APPROVED
MAP-310-07 CN-5310-07 WP-310.3.1 CA-310.3 2026-01-05 baseline map APPROVED
MAP-310-08 CN-5310-08 WP-310.3.2 CA-310.3 2026-01-05 baseline map APPROVED
MAP-310-09 CN-5310-09 WP-310.3.3 CA-310.3 2026-01-05 baseline map APPROVED
MAP-310-10 CN-5310-09 WP-310.3.1 CA-310.3 2026-04-15 MAP CHANGE MC-11 APPROVED
MAP-310-11 CN-5310-10 WP-310.3.4 CA-310.3 2026-01-05 baseline map APPROVED
Abridged — the file continues.
The outcomeWhat a good result looks like
One reconciliation pack in, one row out: which map rows are in force at AS AT, which work packages are closed to this period's charges, which adjustments stand, every control account that does not tie on the cent, every exception and one of five ACR-2026 verdicts. 56 of 64 packs come back with all six graded fields right, against 42 for the best free floor and 0 for the cost tool's own load check. ⚠︎ THAT 56 IS THE STATION RECOMPUTING FROM THE REPLIES' THREE READINGS. The replies' own six fields are whole on 9 of 64 (verdict 23 of 64).
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 8 packs it got wrong are named in the kit README with what it answered, and every one is a READING: 4 map changes approved before AS AT that it left out (CAR-0009, 0016, 0021, 0037), 2 notes dated AFTER AS AT that it followed (CAR-0018, 0031) and 2 reopened packages it kept closed (CAR-0038, 0046). Three of the eight came back TIED on a period the key flags MIS-MAPPED. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your cost system records map change approvals, package reopenings and estimate reversals as a STATUS, and your map extract is refreshed at AS AT — the free modal floor, and do not buy a call at all
42 of 64 packs for $0.00. Unmapped charges, closed packages, mis-loads, double loads and untied accounts are all decidable from columns and dates, and on those 35 packs the floor gets 35 — the same as the paid call. - Map changes, reopenings and reversals arrive as free text — a desk note, an email pasted into the pack — the paid call
This is the whole product. On the 29 packs where a sentence decides a reading the paid arm is 21, the modal floor 7 and the best vocabulary floor 11. - You want a mis-mapped period kept OUT of TIED above all — the paid call, and read the vocabulary floor beside it
A period passed as TIED where the key flags an exception: modal floor 16, negation-skipping floor 7, paid arm 3 (CAR-0009, 0021, 0037 — all three late map-change approvals it left out). The naive vocabulary floor passes only 2 that way, but gets 27 packs whole. - Packs are compiled after AS AT, so notes dated after it are common — the free modal floor for those packs, or strike the notes dated after AS AT in code first
The modal floor reads no note and gets all 7 date traps; the paid arm followed 2 notes dated after AS AT (CAR-0018, CAR-0031). Dropping notes by date is a regex, not a reading. - You want the cost tool's own load check audited — either paid or free — both beat it comprehensively
The load check's status disagrees with ACR-2026 on 36 of 64 packs and gets 0 whole: it returns no reading at all. Its printed status matches the key's verdict on 28 packs. It is published as an arm so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Map changes are routinely approved after the map extract is taken — neither alone — re-export the map at AS AT
The paid arm gets 2 of 6 late approvals, the modal floor 2 and the negation-skipping floor 3: every one of the four notes it missed says the approval came after the actuals were loaded or that the extract has not caught up. A fresh extract makes the reading a column again. - Your programme reconciles to a tolerance, carries multi-currency actuals or reverses actuals across periods — neither, yet
No file in this corpus does. The unit of work is one period's actuals tied on the cent, and every percentage on this page is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own reconciliation packs in the same eight-block shape and data/register.json with your own programme register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per period for arithmetic you already have. That is the case against the best-fitting scenario (“Your cost system records map change approvals, package reopenings and estimate reversals as a STATUS, and your map extract is refreshed at AS AT”). 7 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A cost actual with no charge number — an allocation, an overhead pool, a burden. The whole reduction is the charge-number-to-map-row join; an actual with nothing to join reads as UNMAPPED. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-control-account. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.








