The business caseThe problem this solves
A participation statement has been drafted and somebody has to check it before it goes anywhere near a release review. For each drafted line: which version of the deal's participation clause governs the period, which version the line was actually drafted on, which of the reported receipt records are in that version's base, and which figure each record contributes — because a record routinely prints three or four. The records are printed in the packet, but the facts are in their prose: a reversal booked as a negative entry, a restatement replacing an earlier record, an earned period stated as a duration, a net figure printed before the gross. Today a participation accountant reads them by eye against the standard and re-does the arithmetic. Checking a drafted participation statement against the deal's own clause versions and the reported receipt records behind it, by eye. For every drafted line that means deciding which records are in the governing version's base, which figure each one contributes, what the line recomputes to in cents, and whether the statement may go forward for release review — not the release, not the approval, and not the accountant who owns the statement.
Audience
A participation accountant checking drafted statements, and the reviewer who decides whether one goes forward for release review. The answer this report gives them is NO on the headline: the paid recompute is behind a domain-written free regex on every slice measured, and the kit says so in its first sentence. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual participation statement recompute packets
The corpus is 64 participation statement recompute packets, 0.33 MB (json 4 · jsonl 1 · md 2 · txt 64). Because the shape of a participation failure is not the arithmetic, it is WHICH REPORTED RECORDS are in the governing clause version's base and WHICH FIGURE each one contributes — and both are prose. 423 of the 537 records print more than one figure, 110 print the net before the gross and 92 state no net at all; 82 are reversals booked as negative entries and 39 are restatements replacing an earlier record; 264 dates are long-form, 84 earned periods are stated as a duration and 47 straddle the period boundary; 41 drafted clause versions are implied by a schedule nobody names and 5 version windows close INSIDE the period. 55 of the 122 drafted lines cannot be settled from any column at all. ⚠ v1 OF THIS CORPUS SCORED 95.4 PCT TO A FREE REGEX AND WAS REPLACED BEFORE ANY SPEND — the fix was structural (the clause basis plus five amount phrasings), not a rewording, and data/SOURCES.md item 1 is the full account. A corpus of clean columns would measure a calculator; this one measures a reader, and what it measured is that a domain-written regex reads it better than the paid call.
The corpus
- The 64 participation statement recompute packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 64 statements, data/statement_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your participation statement recompute packets. That is the whole change — there is no database to migrate.
PARTICIPATION STATEMENT — RECOMPUTE PACKET
Marrowfield Pictures, Participations Administration. Prepared for review before release.
This pack recomputes a drafted statement against the deal terms on file and names the
differences. It does not release a participant payment.
PACKET HEADER
Packet id PSR-0001
Title MARROW-5459 "Ashgrove Orchard and the Long Field"
Participant PTY-27197 (an individual participant, represented)
Deal DEAL-8730
Statement period 2025-07-01 through 2025-09-30 (operator-supplied statement cycle)
Statement of record STMT-2025-7076, drafted 2025-10-20
Recompute run RCP-0001
RECOMPUTE SCENARIO
Materiality threshold 500.00
It is an operator-supplied value, not a stated policy, and no policy is cited for it.
DEAL TERMS ON FILE
This operator's own extraction of the deal documents. These are deal terms. No governing
authority is cited for the definition of receipts or for any other term below.
CL-1 Gross participation — theatrical and home video
v1 effective 2023-01-01 through 2024-11-27
basis gross receipts as reported; rate 3.00%; included receipts Home video,
Theatrical exhibition; included territories the domestic territory, the United
Kingdom; deductions distribution fee 30.00%, then off-the-top 5.00%
v2 effective 2024-11-28 onwards
basis gross receipts as reported; rate 4.00%; included receipts Home video,
Theatrical exhibition; included territories the domestic territory, Germany, the
United Kingdom; deductions distribution fee 27.50%, then residual reserve 3.00%
CL-2 Television licence pool participationAbridged — the file continues.
The outcomeWhat a good result looks like
Every drafted participation line carries a finding from a closed list of six, the exact receipt record ids in the governing version's base, the recomputed amount in whole cents and a materiality band; every statement carries its lane (held, ready for release review, or not recomputed), the terms it is held for and its term provenance. A line whose clause terms are missing, whose drafted version disagrees with the governing one, that has already been paid, that rests on a record the base does not support, or that recomputes to a different amount HOLDS the statement.
And when it cannot
⛔ THE PAID RECOMPUTE LOSES TO FREE CODE AND IT IS PUBLISHED AS SUCH. Drafted lines wholly correct — finding, exact receipt ids and recomputed cents together — 59 of 122 after the pure-code station against the domain-written free floor's 74: exact McNemar 28 only-paid against 43 only-floor, p = 0.095924, NOT SIGNIFICANT and the point estimate running AGAINST the call. Headroom above the bar is 48 LINES, never a percentage. As answered, before the station, it is 3 of 122. Swept over all nine slices it leads the best shipped free arm on ZERO and trails on nine — including receipt membership, which was pre-registered as the nearly-free half and where it gets 69 of 122 against the floor's 121 (1 discordant line for it against 53, p < 0.000001 AGAINST). It beats a single fixed reply decisively (52 against 0, p < 0.000001), so it is a real reader; it is just worse than the rules.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Recomputing a drafted participation statement where the analyst's cited column is what you have and the records are printed beside it — the free domain floor
74 of 122 drafted lines wholly correct for $0.00, against the paid recompute's 59 after the station and 3 as answered. It also gets 121 of 122 receipt memberships against the paid arm's 69. - The 55 drafted lines only the reported receipt records settle — no column answers them — the free domain floor, or a person with the records
the paid recompute gets 20 of the 55 against the domain floor's 31 (p = 0.070756, behind) and the tuned ceiling's 55. The 20-against-0 figure that looks like a win is measured against the constant and the columns floor, and both score zero there BY CONSTRUCTION because neither ever opens a receipt record. - A whole-statement sign-off — every line and every statement field right — the free domain floor
33 of 64 statements wholly correct against the paid recompute's 32 after the station (p = 1.000000). A whole statement needs its lane, its held-for list and every line's finding, and on that the two arms are indistinguishable — at $0.00 against $0.041878.
And where nothing here is good enough:
- Deciding which statements reach a release review — neither alone; a person on every statement either arm sends forward
the paid recompute gets the lane right on 48 of 64 statements after the station and the domain floor on 50, p = 0.625 — the corpus cannot separate them, and 0.625 is not a tie, it is an inability to tell. A constant reply through the station reaches 44.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | replace data/corpus/*.txt and data/gold.jsonl with your own statements and key; the rulebook is data/policy.md and the engine reads it as data. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: The paid call on this job. It is behind the floor on all nine slices measured, and on 11 statements the floor takes every drafted line and it takes none. That is the case against the best-fitting scenario (“Recomputing a drafted participation statement where the analyst's cited column is what you have and the records are printed beside it”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A reported receipt record whose earned period is stated in a form the reader does not know. So is a clause version whose effective window the extract does not carry — in both cases the recompute is BLOCKED rather than guessed (R-3). 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Five of the twelve cells this kit measures may never carry a result alone. Both clause-version cells are 100 pct for every arm including a constant reply, because they are read off a printed DEAL TERMS table rather than out of prose; term provenance is the pure-code station's and is 100 pct for every arm; receipt membership leaves the domain floor exactly 1 line of room in 122; and the materiality band is 90.2 pct to the single word NONE. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r002-participation-stmt. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on 2026-09-17 on a copy of the kit folder with the gitignored answer caches (results/cache-*.jsonl) and any .env removed, with no key in the environment: python3 tools/build_corpus.py --check rebuilt all 64 packets, the structured records and the key BYTE-IDENTICALLY in 0.07 s; python3 -m evals.check_labels re-derived the key from its own retyped cascade and record reader with 0 disagreements in 0.10 s; and python3 -m evals.arms re-expressed all four free arms in the paid reading contract and reproduced data/floors.json exactly on all 12 metrics in 0.20 s. python3 -m evals.coverage exits 1 BY DESIGN on the copy as it does here, printing the two rulebook branches this corpus never emits. python3 -m src.app --port 9597 then served the board and every read route answered 200 — the board, the statement list, one statement, the rules, the prompt, the corpus tab and the pressure tab — with the recompute control disabled for want of a key; a deliberate POST to the one route that can spend returned skipped: true and named the four variables to set, and nothing was billed. What the copy could NOT do is buy a call: that needs one credential, and nothing in the free path opens a socket.














