The business caseThe problem this solves
A utility's customer programs office receives rebate applications for efficiency equipment — heat pumps, heat pump water heaters, clothes dryers — and every claimed item has to be checked against the program's own published rules before anybody is paid: was the account active on the install date, is the serial already rebated, is the measure offered to this rate class, was it installed in the program dates and received in time, is the model on the qualified product list, does the premise stay under its cap, does a paid invoice list it. Seven of those are lookups. Three are not: the listing's certified rating for the item's own configuration, the last install date the listing covers, and the rate classes it covers — each stated in a sentence on the qualified product list, and each the place a clerk working through a stack of applications reads the first figure, the first date or the first class word and moves on. Reading every claimed item against the program rules by eye — the account record's service dates, the prior rebates at the premise, the measure table's minimum, cap and offered-to classes, the program dates and the 90-day window, the invoice lines — and then the qualified product listing's three paragraphs that decide the rest. It does not replace the reviewer; every EXCEPTIONS item goes to a person, and nothing here approves, pays, denies or reduces a rebate.
Audience
A program reviewer in a utility's customer programs office working the rebate intake queue before any payment is released. And the program manager who answers for the items that come back. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual rebate applications
The corpus is 64 rebate applications, 0.47 MB (json 4 · jsonl 1 · md 2 · txt 64). Because the shape of a rebate eligibility failure is not the arithmetic, it is the LISTING PROSE. 199 of the 267 items are supported and 20 of the 68 exceptions sit in columns any lookup reads. The 48 that matter are items whose columns are internally perfect and whose listing says something else in a sentence — a rating certified only with a different indoor unit, a last covered date on the other side of 'before' or 'through', a class named only to be excluded — and a corpus without those measures a parser.
The corpus
- The 64 rebate applicationsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 64 applications, data/records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your rebate applications. That is the whole change — there is no database to migrate.
REBATE APPLICATION
APPLICATION HEADER
Application RBT-0001
Utility Quillfeather Basin Electric Cooperative
Program Efficiency Rebates, program year 2026
Received 2026-10-26
Premise PR-64803, 1364 Pellam Road, Marrowby
Intake by D. Farrant, Customer Programs
ACCOUNT RECORD
Account QB-791-7073
Rate class RES (residential)
Service start 2018-12-04
Service end none
Prior rebates at this premise
none
MEASURE TABLE, PROGRAM YEAR 2026
Code Measure Metric Minimum Cap per premise Rebate per unit Offered to
ASHP Ducted air-source heat pump HSPF2 7.8 2 $1,200.00 RES, SMC
MSHP Ductless mini-split heat pump HSPF2 9.0 4 $450.00 RES, SMC, MFC
GSHP Geothermal heat pump COP 3.6 1 $3,000.00 RES
HPWH Heat pump water heater UEF 3.30 1 $600.00 RES, MFC
HPCD Heat pump clothes dryer CEF 5.00 2 $250.00 RES, SMC, MFC
Program installation dates 2026-01-01 through 2026-12-31
Applications received within 90 days of the install date
QUALIFIED PRODUCT LISTINGS
OS-HP673C Ostrin ducted heat pump
Measure ASHP
Rating
With indoor unit OS-C60 the certified rating is HSPF2 8.6.
With indoor unit OS-C36 it is HSPF2 7.8.
Listing dates
Listing active for installations through 2026-12-13.
Account classes
Eligible for residential accounts.
KT-CD335A Kestane heat pump clothes dryer
Measure HPCD
Rating
Certified rating: CEF 5.30.Abridged — the file continues.
The outcomeWhat a good result looks like
Every claimed item carries a finding from a closed list of eleven, the QBEC-REBATE-2026 rule it rests on, and one row copied verbatim as evidence; the application carries NO-EXCEPTIONS or EXCEPTIONS and the exact items a reviewer has to take up. On the published run, after the pure-code station: 261 of 267 item findings, 60 of 64 applications wholly right, 46 of the 48 items only a listing's prose settles.
And when it cannot
⚠︎ THE PAID CALL ON ITS OWN LOSES TO FREE CODE. As the call answered — its own findings, its own answer, its own named items — it gets 202 of 267 items and 22 of 64 applications, where the free column floor gets 219 and 40 for $0.00. It misapplies the lookups: 44 supported items named as exceptions and 12 exceptions called supported, including 0 of the 3 premise-cap items. What wins is src/recheck.py keeping only the call's three listing readings and re-applying the program rules in code. A deployment that took the call's own finding would ship the 202.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your applicants claim correct ratings, and your qualified product list carries end dates and covered classes as structured fields rather than sentences. — the free column floor alone — python3 -m evals.run --floor rules
Every column rule is a lookup, a date comparison or a sum. The floor is 219 of 267 items and 40 of 64 applications on this corpus, deterministic, and $0.00. - Your qualified product list states ratings per configuration, end dates and class exclusions in prose — which real lists do — and applicants copy the headline figure into the Claimed column. — the paid call's three readings, with the rules re-applied in code (src/recheck.py)
Those are the 48 items the column floor is wrong on by construction; the station gets 46 of them and 60 of 64 applications wholly right. - You were going to write keyword readers over the listing prose instead. — read the phrase floor's number first
It is on this page: 139 of 267 items and 3 of 64 applications, worse than reading no prose at all, with 102 supported items named as exceptions.
And where nothing here is good enough:
- You want the pre-check to release payment on NO-EXCEPTIONS without a person looking. — neither, on this evidence
After the station 4 of 64 applications are still wrong, and on RBT-0050 the station turned a clean application into EXCEPTIONS off a date read one day early; the other direction — an ineligible item passed — is 0 today on one run of 64.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own exported applications and data/records.json at your own account, prior-rebate and listing-index records, and replace data/policy.md and data/policy.json with your program's rules. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a reading you do not need, once per application. That is the case against the best-fitting scenario (“Your applicants claim correct ratings, and your qualified product list carries end dates and covered classes as structured fields rather than sentences.”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An application whose panels are not the nine this parser knows. src/rules.py splits on APPLICATION HEADER, ACCOUNT RECORD, MEASURE TABLE, QUALIFIED PRODUCT LISTINGS, CLAIMED ITEMS, INVOICES, APPLICANT NOTES, INTAKE SIGN-OFF and END OF APPLICATION; a missing heading yields an empty section rather than an exception, so a differently shaped application parses to zero items and is scored as zero items. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from a real difference. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-rebate-eligibility. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all three floors score all 64 applications in about a second, and python3 -m evals.check_labels re-derives the whole key from a hand-retyped rulebook. Both are $0.00.






