The business caseThe problem this solves
A statement is printed and beside it sits a folder of supporting schedules. NO SCHEDULE SAYS WHICH FIGURE IT SUPPORTS. Somebody has to decide, figure by figure, what stands behind each number -- and almost none of that is arithmetic. It is a reading problem with an addition at the end of it, and the reading is where it goes wrong. A section subtotal is supported by no paper at all; it is the sum of the three figures printed above it -- and its amount will often ALSO be findable somewhere in the folder, on 63 of this corpus's 240 subtotals, 53 of them with nothing planted, because when a section's figures all come off one omnibus schedule the subtotal simply IS the sum of those lines. A figure supported by four of a schedule's eleven lines is a different answer from the same figure supported by all eleven. And a line carrying a figure's exact amount is evidence of nothing on its own: amounts recur -- prior period comparatives, reversals, group recharges, round contractual sums. Building the tie-out by hand: reading the folder, deciding which schedule is about which caption, picking which of its lines are the figure, adding them up and comparing. Not the review of it, not the sign-off, and not any judgement about whether a figure is right.
Audience
Whoever has to confirm that a printed figure is supported by the file behind it -- the reviewer, the preparer's supervisor, whoever signs off that the tie-out has been done. The output is a tie-out a person CONFIRMS rather than one they build with a highlighter and a calculator. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual statements, each with its whole working-paper folder
The corpus is 60 statements, each with its whole working-paper folder, 0.51 MB (txt 60). The support relationships are the thing, and no public corpus carries them. The plan comes first: tools/build_corpus.py decides what supports what -- whole schedule, four of eleven lines, subtotal, nothing -- and the printed pack is rendered FROM the plan, so nobody looks at a finished pack and decides what supports what, because that decision is exactly what is measured. The case mix is CHOSEN, not sampled: 25 amount collisions planted exact to the cent (12 on subtotals, 13 on detail figures), 27 differences, 22 unsupported figures. And the most interesting collision is not planted at all -- where a section's figures all come off one omnibus schedule, the subtotal is arithmetically equal to a subset of that schedule's lines with no adversary required: 53 of the 63 findable subtotals are this.
The corpus
- The 60 statements, each with its whole working-paper foldergenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your statements, each with its whole working-paper folder. That is the whole change — there is no database to migrate.
STATEMENT TIE-OUT PACK ST-0001
Entity Wraysbury & Hale LLP
Period year ended 31 March 2026
Prepared synthetic. Every entity, figure, schedule and amount is invented.
Amounts printed to two decimal places; tied as whole minor units.
=======================================================================================================
THE STATEMENT AS REPORTED
Ref Caption Amount
-------------------------------------------------------------------------------------------------------
PRACTICE COSTS
F-01 Technical library and reference 10,299.19
F-02 Subcontracted specialist work 6,947.87
F-03 External quality review 8,141.82
F-04 Total practice costs 25,388.88
AMOUNTS RECOVERABLE
F-05 Other debtors 69,573.09
F-06 Trade receivables 32,581.21
F-07 Unbilled work in progress 57,491.84
F-08 Total amounts recoverable 159,646.14
STAFF AND RELATED COSTS
F-09 Salaries - support staff 34,409.26
F-10 Contract and agency staff 38,746.59Abridged — the file continues.
The outcomeWhat a good result looks like
One verdict per captioned figure, with its citation, under a standard written down in full and given to every arm. On the 60-statement corpus, rechecked: 842 of 844 figures correct on all three graded things at once (99.8 pct), 58 of 60 statements entirely correct, 555 of 555 TIED figures with the right lines, 27 of 27 differences reported and 26 of them exact to the minor unit, and all 25 planted amount collisions resolved with the decoy line cited zero times.
And when it cannot
A figure called TIED against the wrong lines. That is the expensive kind of right answer: the verdict is what everyone reads, it says the work has been done, and nobody opens the schedule again. It is why this kit scores the SUPPORT SET and not only the verdict -- and it is exactly where the free floor fails, tying 63 subtotals to schedules that have nothing to do with them and citing the planted decoy line on 25 of 25 collisions.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Statements whose subtotals are captioned predictably and whose folders hold no repeated amounts — free code -- and give it the caption rule
Measured here: the caption rule plus an exact-sum walk takes the free floor to 821 of 844 figures (97.3 pct), 41 of 60 statements and 21 of 21 clean statements, for $0.00 and 0.3 s. On a collision-free folder, rung 1's exhaustive subset enumeration is already 546 of 555 on TIED support and 349 of 353 on whole-schedule figures. - Folders where amounts recur -- comparatives, reversals, recharges, round contractual sums -- so a line carrying a figure's exact amount may have nothing to do with it — the model, rechecked
This is the whole case and it is what survives every attack on the headline. 25 of 25 planted collisions resolved with the decoy cited 0 times, against the shipped floor's 0 of 25 with the decoy cited 25 of 25, and against the caption-rule floor's 12 of 25 with the decoy cited 13. 13 of the model's 22 remaining wins over the best free arm are planted collisions and 8 more are subset DIFFERENCEs. Beside it: 27 of 27 differences reported against 14, 0 false ties against 67, and 0 missed differences against 4.
And where nothing here is good enough:
- Any fork that keeps the model but drops src/recheck.py — do not
The RAW arm is 502 of 844 and LOSES to the free floor's 581. 340 of its 342 misses are TIE-1.6's whole-schedule shorthand and would ship as wrong citations; the recheck's 541 overrides are what take the arm to 842. A fork that skips it pays for the call and publishes the worse of two answers.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own packs as .txt into data/corpus/ in the printed shape the generator emits -- a figure table of F-nn caption amount, then SCH-<letter> blocks of <letter>-nn description amount with a SCHEDULE <letter> TOTAL line -- and write data/gold.jsonl with, per figure, the verdict, the schedule, the lines and (for a subtotal) derived_from. ⚠︎ THE WHOLE WORKING-PAPER FOLDER REACHES YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying $0.0075 a statement and a 70.6 s p50 for a job that an exact-sum search and a string test already do -- and inheriting a 212.5 s worst call to do it. That is the case against the best-fitting scenario (“Statements whose subtotals are captioned predictably and whose folders hold no repeated amounts”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | Every statement has exactly four subtotals -- three section subtotals and a grand total, on all 60, without exception. Real statements have one level of subtotalling or four, nested, with cross-casts and carried-forward pages. 9 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | The adversarial arm has NEVER BEEN FIRED. evals/injection.py ships written, paired and unrun; there is no x001 result, no suppression rate and no citation-movement count for this kit. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-figure-tieout. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — ⚠︎ THE CORPUS DOES NOT REBUILD, AND THIS KIT'S CLAIM WAS TESTED RATHER THAN REPEATED. python3 tools/build_corpus.py --check FAILS. Measured in this capture pass at PYTHONHASHSEED 0, 1, 7 and 97531: SIX of the 60 statements differ from the committed bytes at three of those seeds and five at the fourth -- always the same six (ST-0014, ST-0020, ST-0023, ST-0038, ST-0046, ST-0054), with ST-0038 matching by luck at seed 7. The cause is collide_line = [r for r in placed ...] at tools/build_corpus.py:389, where placed is a SET, so which figures receive a planted collision depends on str-hash randomisation. ⚠︎ AND THE DIVERGENCE IS LARGER THAN THE KIT'S OWN PROSE SAYS. README and data/SOURCES.md both describe it as 'always in the decoy lines of the noise schedule'; measured, data/gold.jsonl ALSO differs on those same six statements -- plants.collisions[].figure and the matching answer.figures[].collision_decoy MOVE TO A DIFFERENT FIGURE (F-05 to F-02, F-05 to F-11, and so on). A rebuild therefore plants its 25 collisions on a different 25 figures than the ones the published 25/25 was measured on. Two further facts, both measured: (a) --check compares ONLY data/corpus/*.txt, so it does not see the gold.jsonl divergence and under-reports its own failure; (b) data/corpus-stats.json is also not reproducible, differing by one byte in the bytes field (538,453 rebuilt against 538,454 committed). WHAT IS NOT DAMAGED, and this was checked rather than assumed: every published COUNT is identical across all four seeds (844 figures, 240/555/27/22, 25 collisions, 39 hard, the eleven case counts), the committed corpus passes all 15 of evals/check_labels.py's checks, and it is exactly what every run on disk measured -- so the kit's NUMBERS stand and its REPRODUCIBILITY claim does not. The one-line fix (sorted(placed)) still diverges from the committed bytes, so applying it means a corpus rebuild and a re-bought paid run: a spend decision, deliberately left to the operator. The free floor WAS re-derived here from source with zero calls and reproduced exactly -- 581/844, 0/240 SELF-DERIVED, 67 false ties, 25 decoys cited -- and the paid arm was re-scored from results/cache-r001-figure-tieout.jsonl, also with zero calls.



