The business caseThe problem this solves
Somebody holds an application in one hand and a folder of documents in the other. The application states four to six values, each over its OWN window; the folder holds dated records — wage statements, award letters, a self-employment ledger, an account statement, a third-party verification record — and the difficulty is almost never finding the numbers. It is that a record covers 1–15 when the item covers 1–31, that a payroll agent's name sits where the employer's should be, that a corrected statement replaces one filed a fortnight earlier, and that a one-time adjustment is named in a record's own text rather than in its amount column. Every one of those turns a comparison into a judgement about what is even being compared. Building the field-by-field comparison by hand — deciding which record speaks to which stated item across differently written names and trading names, noticing which records a correction replaced, adding what is left, and checking the window is actually covered. It replaces the BUILDING. It does not replace the adjudicating, and there is no field in the answer contract that could.
Audience
An eligibility worker who is going to adjudicate this list, and whoever decides whether a pack like it is worth running at all. Its own answer on the headline is a qualified no: a free rules floor plus a free pure-code station reaches 96.2 pct of items and a paid call reaches 98.5 pct, and a paired exact test cannot separate them at 0.05. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual verification packet
The corpus is 60 verification packet, 0.13 MB (txt 60). It had to be able to separate arms, and the FIRST BUILD COULD NOT. An independent regex parser with its own retyped rulebook (evals/check_labels.py) recovered 252 of 252 item verdicts from the packet text alone — a corpus that measures nothing. Three families were added whose deciding fact is in a record's THIRD LINE rather than in either printed table: payroll_agent, superseded_record and award_adjustment. That checker now recovers 252 of 262 (96.2 pct), and the ten items it cannot reach are exactly the ten a paid call is being asked to earn its money on.
The corpus
- The 60 verification packetgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states what is invented (all of it), what the generator costs the measurement, and the five things a synthetic key cannot settle.
Swap this folder for your own material and the kit is pointed at your verification packet. That is the whole change — there is no database to migrate.
BENEFITS ELIGIBILITY VERIFICATION PACKET BEV-0001
Prepared 2026-09-02 under VCS-2026 | verification window 2026-08-01 to 2026-08-31
CASE FACTS
Case reference BEV-0001
Verification window 2026-08-01 to 2026-08-31
Resource test date 2026-08-04
Household Tomas Petrov (applicant)
STATED ON THE APPLICATION
item person what is stated stated value window
S1 Tomas Petrov wages from Northside Grocery Co. $2,005.00 2026-08-01 to 2026-08-31
S2 Tomas Petrov self-employment net income, Alder Lane Alterations $624.00 2026-08-01 to 2026-08-31
S3 Tomas Petrov monthly payment from Coastal Mutual Disability Plan $327.00 2026-08-01 to 2026-08-31
S4 household countable resources held at Pinehurst Mutual Savings $1,579.00 2026-08-04 to 2026-08-04
EVIDENCE FILED
E1 WAGE-STATEMENT covers 2026-08-01 to 2026-08-15
names Tomas Petrov | payer Northside Grocery Co. | amount for this period $1,002.50
E2 AWARD-LETTER covers 2026-08-01 to 2026-08-31
names T. Petrov | source Coastal Mutual Disability Plan | amount for this period $327.00
a one-time adjustment of $82.00 was included in this period's payment
E3 ACCOUNT-STATEMENT covers 2026-08-01 to 2026-08-31
names Tomas Petrov | institution Pinehurst Mutual Savings | balance on the test date $1,579.00
CASE NOTES
The household was interviewed on 2026-09-11 and confirmed the composition above; no change was made to what is stated.
The outcomeWhat a good result looks like
A comparison a person CONFIRMS instead of one they build: every stated item reported with a verdict, the E-numbers behind it and what those records sum to — so the arithmetic can be checked rather than believed, and the worker adjudicates from there.
And when it cannot
It asserts a conflict a household does not owe. That is the failure this kit counts by name and splits four ways, and the biggest bucket is a SHORT WINDOW filled in by arithmetic: the free floor did it 13 times before the pure-code station and 0 after; the paid call did it 0 times.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your packets are a fixed layout and every record's coverage dates are printed — the free rules floor plus src/recheck.py — 96.2 pct of items, $0.00
the paid call reaches 98.5 pct and a paired exact test cannot separate them (p = 0.070). Six items over 262 is not a purchase. - Corrections, payroll agents and adjustments are common in your file — the paid call
all four of the floor's surviving false conflicts and three more of its misses are exactly those three shapes, and the call got every one of them.
And where nothing here is good enough:
- You need to be sure a conflict is not MISSED — neither, yet
conflict recall is 12 of 15 on both arms, and all three misses are one mechanism (an adjustment named in a record's own text). Fix the reading before choosing an arm. - Your evidence coverage dates have to be read off a document — nothing here, measured
data/cases.json hands both arms every record's class and dates, so the whole coverage half of this job is arithmetic in this kit. That case is not measured here at all.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/ with your own packets and data/cases.json with your own evidence index (per record: source class, coverage start, coverage end), then write your own gold.jsonl or derive it the way tools/build_corpus.py does. ⚠︎ DO NOT BRING REAL CASE FILES. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying a call to do arithmetic. The station supplies VC-3 and VC-6 for nothing. That is the case against the best-fitting scenario (“Your packets are a fixed layout and every record's coverage dates are printed”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet that is not this layout. Both free arms and the independent checker parse two fixed-width tables and two lines per evidence record; a scan, a spreadsheet export or free prose moves every number on this page first and furthest. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A second run of the same corpus. One run, and the headline margin is six items out of 262. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-02 — r001-stated-documented. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clone with no key and no network regenerates the corpus, scores both free floors over all 262 items, replays every recorded run and renders the whole board — measured at 199 ms for the generate-and-score half. Only --yes without --floor and the board's one POST route reach a provider.



