Home › Use Cases › Compare what a household stated against what its documents show, and list the conflicts
Use caseUC0282
🧪 Use-case kit · runnable

Compare what a household stated against what its documents show, and list the conflicts

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Somebody holds an application in one hand and a folder of documents in the other. The application states four to six values, each over its OWN window; the folder holds dated records — wage statements, award letters, a self-employment ledger, an account statement, a third-party verification record — and the difficulty is almost never finding the numbers. It is that a record covers 1–15 when the item covers 1–31, that a payroll agent's name sits where the employer's should be, that a corrected statement replaces one filed a fortnight earlier, and that a one-time adjustment is named in a record's own text rather than in its amount column. Every one of those turns a comparison into a judgement about what is even being compared. Building the field-by-field comparison by hand — deciding which record speaks to which stated item across differently written names and trading names, noticing which records a correction replaced, adding what is left, and checking the window is actually covered. It replaces the BUILDING. It does not replace the adjudicating, and there is no field in the answer contract that could.

Audience

An eligibility worker who is going to adjudicate this list, and whoever decides whether a pack like it is worth running at all. Its own answer on the headline is a qualified no: a free rules floor plus a free pure-code station reaches 96.2 pct of items and a paid call reaches 98.5 pct, and a paired exact test cannot separate them at 0.05. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual verification packet

The corpus is 60 verification packet, 0.13 MB (txt 60). It had to be able to separate arms, and the FIRST BUILD COULD NOT. An independent regex parser with its own retyped rulebook (evals/check_labels.py) recovered 252 of 252 item verdicts from the packet text alone — a corpus that measures nothing. Three families were added whose deciding fact is in a record's THIRD LINE rather than in either printed table: payroll_agent, superseded_record and award_adjustment. That checker now recovers 252 of 262 (96.2 pct), and the ten items it cannot reach are exactly the ten a paid call is being asked to earn its money on.

The corpus

  • The 60 verification packetgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states what is invented (all of it), what the generator costs the measurement, and the five things a synthetic key cannot settle.

Swap this folder for your own material and the kit is pointed at your verification packet. That is the whole change — there is no database to migrate.

One verification packet, as the model receives itBEV-0001.txt · 1 of 60
BENEFITS ELIGIBILITY VERIFICATION PACKET                              BEV-0001
Prepared 2026-09-02 under VCS-2026 | verification window 2026-08-01 to 2026-08-31

CASE FACTS
  Case reference               BEV-0001
  Verification window          2026-08-01 to 2026-08-31
  Resource test date           2026-08-04
  Household                    Tomas Petrov (applicant)

STATED ON THE APPLICATION
  item  person              what is stated                                             stated value  window
  S1    Tomas Petrov        wages from Northside Grocery Co.                             $2,005.00  2026-08-01 to 2026-08-31
  S2    Tomas Petrov        self-employment net income, Alder Lane Alterations             $624.00  2026-08-01 to 2026-08-31
  S3    Tomas Petrov        monthly payment from Coastal Mutual Disability Plan            $327.00  2026-08-01 to 2026-08-31
  S4    household           countable resources held at Pinehurst Mutual Savings         $1,579.00  2026-08-04 to 2026-08-04

EVIDENCE FILED
  E1   WAGE-STATEMENT              covers 2026-08-01 to 2026-08-15
       names Tomas Petrov           | payer       Northside Grocery Co.              | amount for this period $1,002.50
  E2   AWARD-LETTER                covers 2026-08-01 to 2026-08-31
       names T. Petrov              | source      Coastal Mutual Disability Plan     | amount for this period $327.00
       a one-time adjustment of $82.00 was included in this period's payment
  E3   ACCOUNT-STATEMENT           covers 2026-08-01 to 2026-08-31
       names Tomas Petrov           | institution Pinehurst Mutual Savings           | balance on the test date $1,579.00

CASE NOTES
  The household was interviewed on 2026-09-11 and confirmed the composition above; no change was made to what is stated.

The outcomeWhat a good result looks like

A comparison a person CONFIRMS instead of one they build: every stated item reported with a verdict, the E-numbers behind it and what those records sum to — so the arithmetic can be checked rather than believed, and the worker adjudicates from there.

And when it cannot

It asserts a conflict a household does not owe. That is the failure this kit counts by name and splits four ways, and the biggest bucket is a SHORT WINDOW filled in by arithmetic: the free floor did it 13 times before the pure-code station and 0 after; the paid call did it 0 times.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your packets are a fixed layout and every record's coverage dates are printed — the free rules floor plus src/recheck.py — 96.2 pct of items, $0.00
    the paid call reaches 98.5 pct and a paired exact test cannot separate them (p = 0.070). Six items over 262 is not a purchase.
  • Corrections, payroll agents and adjustments are common in your file — the paid call
    all four of the floor's surviving false conflicts and three more of its misses are exactly those three shapes, and the call got every one of them.

And where nothing here is good enough:

  • You need to be sure a conflict is not MISSED — neither, yet
    conflict recall is 12 of 15 on both arms, and all three misses are one mechanism (an adjustment named in a record's own text). Fix the reading before choosing an arm.
  • Your evidence coverage dates have to be read off a document — nothing here, measured
    data/cases.json hands both arms every record's class and dates, so the whole coverage half of this job is arithmetic in this kit. That case is not measured here at all.

At a glanceHow the whole thing runs

98%item verdict pct
69,791 msp50, end to end
$30.11per 1,000 stated items · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/ with your own packets and data/cases.json with your own evidence index (per record: source class, coverage start, coverage end), then write your own gold.jsonl or derive it the way tools/build_corpus.py does. ⚠︎ DO NOT BRING REAL CASE FILES. Corpus lens →
When is this the wrong choice?Avoid: Buying a call to do arithmetic. The station supplies VC-3 and VC-6 for nothing. That is the case against the best-fitting scenario (“Your packets are a fixed layout and every record's coverage dates are printed”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet that is not this layout. Both free arms and the independent checker parse two fixed-width tables and two lines per evidence record; a scan, a spreadsheet export or free prose moves every number on this page first and furthest. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second run of the same corpus. One run, and the headline margin is six items out of 262. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-stated-documented. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clone with no key and no network regenerates the corpus, scores both free floors over all 262 items, replays every recorded run and renders the whole board — measured at 199 ms for the generate-and-score half. Only --yes without --floor and the board's one POST route reach a provider.

A living map of modern AI — kept current every morning