Home › Use Cases › Recompute each month's benefit under the rule version that governed it
Use caseUC0457
🧪 Use-case kit · runnable

Recompute each month's benefit under the rule version that governed it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A county reopens a household's review period and has to say, month by month, what the benefit SHOULD have been — under the rule version that governed THAT month, not the one in force today. The facts are on the worksheet, but two of them are not: which amendment notice actually reached each benefit month, and whether the household's shelter cap was lifted. Today a worker reads the notices and the case narrative by eye against a printed calculation standard, and a second worker re-reads the ones that produce a difference. Reading every benefit month of a reopened review period against the amendment notices and the case narrative by eye, then running the ten-step calculation by hand for each one — not the decision that follows, and not the worker who owns it.

Audience

A benefit recomputation reviewer reopening one household's review period, and the supervisor who decides whether the result goes any further. The answer this report gives them is a qualified no: on the headline measure a free keyword pass over the same notices beats the paid call, and the kit says so in its first sentence. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual benefit recomputation workpapers

The corpus is 64 benefit recomputation workpapers, 0.49 MB (json 3 · jsonl 1 · md 2 · txt 64). Because the shape of a benefit recomputation failure is not the arithmetic, it is WHICH RULE VERSION the arithmetic runs under, and that is prose. Every workpaper carries a 'Current rule version v4' line in its header — applying it uniformly is the cheapest wrong answer, and 9 workpapers' own VerUsed column agrees with it. 11 amendment notices were WITHDRAWN before they took effect. One sentence distinguishes determinations made on or after a date from benefit months beginning on or after it. The adoption date and the first governed month are two and a half months apart, and the allowance gap between adjacent versions is 7 to 25 dollars, which reads as rounding. A corpus of clean arithmetic would measure a calculator; this one measures a reader.

The corpus

  • The 64 benefit recomputation workpapersgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 64 workpapers, data/case_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your benefit recomputation workpapers. That is the whole change — there is no database to migrate.

One benefit recomputation workpaper, as the model receives itBRC-0001.txt · 1 of 64
BENEFIT RECOMPUTATION WORKPAPER

WORKPAPER HEADER
  Workpaper            BRC-0001
  Jurisdiction         Marrowdale County
  Programme            Household Support Allowance
  Calculation rules    HSA-CALC-2026
  Case reference       MC-40013
  Review period        2025-09 through 2026-02 (6 benefit months)
  Current rule version v4
  Finding recorded     2026-01-05
  Workpaper dated      2026-01-11
  Prepared by          R. Okonjo, Household Assistance Unit

DECLARED COUNTY STANDARDS
  Supervisor review threshold   250 dollars, absolute net difference for the review period
  Calculation timeliness        30 calendar days from the finding date
  Elderly age                   60 years
  These are Marrowdale County's own declared figures. No regulation is cited for any of them.

RULE VERSION FIGURES
  Version  Allow1  Allow2  Allow3  Allow4  Allow5  StdDed  EarnDis%  ShelThr%  ShelCap  RedRate%  MinBen  MinSizes
  v1          291     535     766     973    1155     198        20        50      672        30      23  1, 2
  v2          298     546     782     993    1179     204        20        50      697        30      23  1, 2
  v3          305     558     799    1015    1205     209        22        50      712        30      24  1, 2
  v4          313     573     820    1042    1237     215        22        48      744        28      25  1, 2

RULE AMENDMENT NOTICES
  N-1   Beginning with the benefit month of January 2024, version v1 governs.
  N-2   Beginning with the benefit month of October 2024, version v2 governs.
  N-3   Beginning with the benefit month of April 2025, version v3 governs.

Abridged — the file continues.

The outcomeWhat a good result looks like

Every benefit month carries a finding from a closed list of five, the recomputed amount under the rule version the notices settle for it, the difference against what was issued, and one line quoted verbatim from the workpaper as evidence — plus a review route, a net difference, a flagged-month count, a completeness flag and a timeliness verdict for the review period as a whole. A month whose figures are missing or whose rule version the notices do not settle is FLAGGED and never estimated.

And when it cannot

⚠︎ THE HEADLINE IS A LOSS AND IT IS PUBLISHED AS ONE. The measure is the RECOMPUTED AMOUNT, and the paid call takes 304 of 384 against the best FREE arm's 342 — a keyword reader over the same amendment notices, at $0.00 and no network. p = 0.000077 on the paired exact test, 26 months won against 64 lost. Worse for the finding rate: a CONSTANT reply that answers AS-ISSUED everywhere and echoes the Issued column scores 268 amounts and BEATS a real pure-code calculator running all ten steps of the standard (266). On the 118 months whose answer lives only in the prose the paid call and that constant reply are NOT SEPARABLE (77 v 75, p = 0.86). What the call does buy is nameable and narrow: 44 of the 55 months the worksheet coded to the wrong rule version, where every column-only floor takes 0.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A reopened review period whose amendment notices moved the governing rule version on some months and whose worksheet coding did not follow — the paid call, with the recheck station
    44 of the 55 months the worksheet coded wrong, against 0 for every column-only floor. This is the only row on this page where the money is clearly buying something no free arm reaches.
  • A review period whose months must be DECLINED — no live notice reaches them, or two reach them with different versions — the free keyword floor
    18 of 29 against the paid call's 6. Declining is this kit's worst measured behaviour and the free arm is three times better at it.
  • A first pass that decides which months a worker looks at — the free columns calculator beside the paid call, and read the disagreement
    the columns floor is right on all 266 column-settled months and the paid call is right on the 55 the columns misstate; the set where they disagree is exactly the set worth a person's time.

And where nothing here is good enough:

  • A whole-workpaper sign-off — every month and all five period totals right — neither; a person
    0 of 64 workpapers are wholly correct on the paid call and 46 of 64 on the best free arm. A workpaper needs everything right at once and the paid arm never gets there on this corpus.

At a glanceHow the whole thing runs

79%recomputed amount correct pct
3,061 msp50, end to end
$1.35per 1,000 benefit recomputation workpapers · GPT-5.6 Luna

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own rendered workpapers and data/case_records.json at your own structured half, then rewrite src/workpaper.py for your layout and data/policy.md + data/policy.json for your calculation standard. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: If your notices are formulaic. A keyword floor over the same sentences takes 342 of 384 amounts here against the paid call's 304, for $0.00 — write the patterns first and only buy a call if they fail. That is the case against the best-fitting scenario (“A reopened review period whose amendment notices moved the governing rule version on some months and whose worksheet coding did not follow”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A review period longer than about a dozen benefit months. The measured worst case is 7 months and 20 of the 64 workpapers are that long; the largest reply drew 868 of the 1,000-token ceiling, so roughly twice this length truncates. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real county's workpapers. Every figure here is against a key derived from a generator, on an invented calculation standard, for a programme that does not exist. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-benefit-recompute. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on a clean checkout of this repository with no key configured and no network: python3 -m evals.baseline scores all four free floors over all 64 workpapers and 384 benefit months and writes the four b000 result files, and python3 -m src.app serves the whole board — every panel, the committed run replayed, and all four floors computed live — with the button that would call the model disabled and the reason printed beside it. What a clean checkout CANNOT do is buy a call: that needs one credential in .env, and nothing in the free path opens a socket.

A living map of modern AI — kept current every morning