Decide a car dealer's warranty claim, line by line
A dealer sends in a repair bill with parts, labour and a story behind each line. This app checks every line against the policy, the labour guide and the correspondence, and tells you what to pay.
PresenterOpens the private repo. Visible to admins only.
For the warranty analystAutomotive
Why it matters
Today's manual process, and the same job with the app
A warranty analyst at a vehicle manufacturer, deciding whether to pay or push back a dealer's repair bill.
✕Today's manual process
1Read the claim line by line against the coverage terms and the labour time guide.
2Look up every operation in the labour guide, and compare the hours claimed to what it publishes.
3Search the correspondence for the one sentence that changes the answer: a coverage note, a hotline case, a prior repair.
4One line missed means paying a claim that should have been denied, or denying one that should have been paid.
Every claim read manually, line by line
✓With the app
1The claim is read against the policy, the labour guide and every record on file, automatically.
2Every operation is checked hours claimed against hours published, on every line, whether or not anything looks wrong.
3The correspondence is searched for the sentence that reopens a coverage window or changes the payable time.
4Every line gets a reason payable, denied with the record behind it, or flagged as not enough to go on.
Every line reasoned, with its record
See it work
One real case: what the app found, step by step
WC-0018: a 2023 Calder Ridgeline, 14 months in service, with six claimed lines and one coverage note that changes the answer.
Decide a car dealer's warranty claim, line by lineReference appBuilt to be shaped to your process
5
1What was claimed The alternator: 1.60 hours claimed, exactly what the guide publishes. Payable.
2Why, beyond the tables Cites the coverage note that reopened this operation past its printed mileage limit.
3Checked, not hidden The page flags that citation itself, since it is not a document record number.
4A line it denies Drive shaft hours run above the guide's published time.
5Proof, not just a verdict A prior repair to the same part is on file, with no sign-off.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Decide a car dealer's warranty claim, line by line
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A dealer has transmitted a warranty claim and someone has to decide whether it is payable and at what labour time before the money moves. That decision is not a date subtraction -- it is a DENIAL YOU WILL HAVE TO DEFEND WHEN THE DEALER APPEALS IT, and one you may have to reverse in front of a bulletin the manufacturer issued itself. So the failure that costs is not a wrong coverage window. It is a line denied on a window the pack has already reopened: a special coverage adjustment that extended that operation on that model, a technical hotline case that authorised the extra diagnostic time, a prior repair carried out under the district manager's own goodwill authorisation and held at the region rather than stapled to the repair order. The coverage really is closed in all three and the denial is still wrong. And the other direction costs the manufacturer instead: a combination operation, where the hours claimed equal the published time exactly and the published time is not the time payable, reads as perfect line by line. Over-payment is fraud exposure; wrong denial is a dealer relations failure, and on this corpus every arm makes one of them. A warranty analyst reading a transmitted claim line by line against the policy: working out completed months from the in-service date to the repair date and comparing the odometer against the limit for whichever coverage governs that operation, looking each operation code up in the labour time guide and comparing the hours claimed against the published time, classifying each cause code as a covered defect or as wear, maintenance or customer damage, checking the file for a prior repair to the same component that nobody authorised, and then -- the part no table helps with -- reading the service correspondence for the one sentence that reverses any of it.
Audience
Warranty analysts and claim auditors at a vehicle manufacturer who adjudicate transmitted dealer claims before payment or chargeback, the zone and district staff who answer the appeals, and the people who build tooling for either. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual warranty claim review packs
The corpus is 45 warranty claim review packs, 0.43 MB (txt 45). A dealer warranty claim carries a VIN, a named customer, a dealer code, a repair order and a labour rate that is commercially confidential between a manufacturer and its network, and the policy and labour time guide it is adjudicated against are the manufacturer's own. There is no public set of (claim, adjudication) pairs and there is not going to be one, and publishing a scrubbed real set would be worse than publishing none, because the scrubbing is precisely where the interesting defect hides -- the sentence naming a special coverage adjustment is the first thing a scrubber removes and the only thing that makes the case a case. The harder reason is that the thing being measured has to be PLANTED to be measured at all. The question this kit asks is not "is this claim outside its coverage" but "is this a denial you could defend when the dealer appeals it holding the manufacturer's own document", and a real archive does not come labelled with which apparent denial grounds had already been reopened by a bulletin, a hotline case or a district manager. So the corpus is generated, the generator ships, the seed is recorded, and the answer key is re-derived by a gate rather than trusted.
The corpus
The 45 warranty claim review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your warranty claim review packs. That is the whole change — there is no database to migrate.
One warranty claim review pack, as the model receives itWC-0001.txt · 1 of 45
Warranty Claim Review
----------------------------------------------------------------
Claim WC-0001
Dealer Harrowgate Motor Group, dealer code D-6613
Vehicle 2024 Calder Ridgeline Hybrid
VIN 5JVXEB6FSXW7PVWD0
Repair order RO-252396
Repair date 2026-02-28
In service date 2023-12-28
Odometer at repair 23,099
Labour rate 134.50
Claim basis published labour operation times, coverage under the vehicle warranty policy
Warranty Policy And Guides
----------------------------------------------------------------
Coverage terms
CV-BAS Bumper to bumper, basic vehicle 36 months 36,000 miles WP-2.1
CV-PWR Powertrain, engine transmission and drive 60 months 60,000 miles WP-2.2
CV-EMI Emissions related components 96 months 80,000 miles WP-2.3
Labour operation guide
LO-1120 Alternator, remove and replace CV-BAS 1.60 hours WP-3.1
LO-1210 Door mirror assembly, replace CV-BAS 0.70 hours WP-3.1
LO-1408 Air conditioning compressor, replace CV-BAS 3.40 hours WP-3.1
LO-1470 Seat frame, front, remove and replace CV-BAS 2.80 hours WP-3.1
LO-1512 Infotainment head unit, replace CV-BAS 1.40 hours WP-3.1
LO-1580 Tailgate strut pair, replace CV-BAS 0.60 hours WP-3.1
LO-2210 Cylinder head gasket, replace CV-PWR 11.60 hours WP-3.1
LO-2255 Rocker cover gasket, replace CV-PWR 1.10 hours WP-3.1
LO-2404 Transmission valve body, replace CV-PWR 5.10 hours WP-3.1
Abridged — the file continues.
The outcomeWhat a good result looks like
A drafted adjudication: one entry per claimed line, in the pack's own order, each carrying DENY with one of six named grounds, the policy provision from this pack and the record identifiers behind it -- or PAYABLE, or INSUFFICIENT_EVIDENCE naming what is missing -- plus the labour time the operation guide publishes, read independently on EVERY line whether or not anything is wrong with it, and a claim-level recommendation. Beside every line, what pure code alone would have written, which on this corpus scores HIGHER than the paid arm.
And when it cannot
⚠︎ THE STRONGEST FREE FLOOR BEATS THE PAID ARM ON THE DISCRIMINATOR AND THIS PAGE LEADS WITH IT. 76.30 pct against 44.44 pct over the same 270 claimed lines (b002-warranty-claim-claimgate against r001-warranty-claim), free, offline, no key, and 0.1 s of wall clock for all 45 packs against 1,503.4 s. The 31.86-point gap is three measured things and none of them is the arm failing to find the defect. ALL 62 of its over-denials are one shape: the dealer's own rounding, 0.01 h on 23 lines, 0.02 h on 13 and 0.03 h on 26, above the published operation and denied as LABOUR_TIME_EXCEEDED -- not one over-denial is anything else. The free floor's fix is one float, src/checks.TOL = 0.05 hours; the arm has no constant to set. Second, it cited WP-3.4 on all 42 labour-time cells where the key carries WP-3.1, and WP-3.4 is arguably the better clause -- THE MODEL CONVICTED THE ANSWER KEY, THE KEY WAS NOT CHANGED AND THE RUN WAS NOT RE-FIRED. Third, 72 of 269 citing lines quoted a real string out of the pack -- an operation code, a hotline case, a goodwill authorisation, a bulletin number -- into a field that asked for a record identifier an analyst can pull; 0 of them were absent from the pack. Credit it both arguable readings and it reaches 70.74 pct, still 5.56 points behind free code.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Every fact that decides a claim on your bench is already a FIELD -- an in-service date, an odometer reading, a cause code in a table with a covered/not-covered column, an operation code in a published time guide, a prior-repair record with an authorisation number on it -- and your dealers claim labour in hundredths of an hour. — the free floor -- claim-gate, $0.00, and set TOL to your own rounding It scores 76.30 pct on the discriminator against the fast tier's 44.44, cites the provision correctly on 100 pct of its denials against 66.15, names all 16 unsettleable lines against 6, reads the published labour time right on 261 of 261 against 239, and invents nothing. It runs all 45 packs in 0.1 s with no key and no network.
Your bulletins, hotline authorisations and goodwill decisions live in prose -- a technical information system, a CRM case, a region's spreadsheet -- and a claim that looks out of coverage on the printed policy is routinely reopened by one of them. — the paid arm, and budget for the over-denials This is the only column free code does not reach at any tolerance: 22 of 22 for the fast tier against 0 of 22 for all three floors, and the ablation proves it is real reading -- blind the same arm and it drops to 0 of 22 and lands on the floor's figure in five columns. It also cuts the expensive failure by three and a half times: 12 silent over-denials of 42 against the floor's 42 of 42.
You are pointing this at real transmitted claim records out of a dealer management system. — rewrite src/claim.py first, and expect the correspondence column to disappear src/claim.py parses this corpus's fixed-column layout and returns EMPTY tables against anything else rather than guessing, so every line would come back INSUFFICIENT_EVIDENCE -- an honest failure, not a hidden one. And the half this kit measures as the paid arm's premium is exactly the half that is not in the claim feed at all in a real estate. data/SOURCES.md says what breaks, field by field.
And where nothing here is good enough:
You want one number to decide with. — neither -- read the column split ⚠︎ THE SINGLE NUMBER IS THE MOST MISLEADING THING ON THIS PAGE. The assigned floor ties every arm on the date column (20 of 20) and the mileage column (18 of 18) and scores 0 of 34 and 0 of 16 on the two cause columns. The strongest floor closes both cause columns for nothing and still takes 0 of 22 on the combination-operation half of labour time. The paid arm takes that 22 and pays for it with 62 over-denials. No single accuracy figure carries any of that.
Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Decide a car dealer's warranty claim, line by line14 steps · 4 questions · run once, for real · 2026-08-26
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
tools/build_corpus.py writes data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json. src/select.py withholds a NAMED SECTION and is not a redaction system, and this corpus contains its own counter-example.Corpus lens →
When is this the wrong choice?
Avoid: ⚠︎ Do not use the free floor where any decisive fact is a sentence. It takes 0 of 22 correspondence grounds and denies all 42 of the lines the pack itself has already answered -- sent as written, that is 42 denials to dealers holding the manufacturer's own bulletin, hotline case or goodwill authorisation. And do not take its 76.30 pct on trust: it reads 57.78 pct at TOL = 0.00, so sweep the constant on your own corpus before believing any of it. That is the case against the best-fitting scenario (“Every fact that decides a claim on your bench is already a FIELD -- an in-service date, an odometer reading, a cause code in a table with a covered/not-covered column, an operation code in a published time guide, a prior-repair record with an authorisation number on it -- and your dealers claim labour in hundredths of an hour.”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A claim that is not this shape. src/segment.py looks for underlined headings and src/claim.py is a set of regular expressions written for these fixed-width tables and these key/value blocks. 9 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
The 72 lines carrying a non-identifier citation are scored as wrong and there is a real argument they are half-right: 0 of them are absent from the pack, so the arm quoted a genuine bulletin, hotline case or operation code into a field that wanted a record identifier. Both the strict figure (26.77 pct) and the absent-from-pack figure (0.00 pct) are published so a reader can choose which they believe. 10 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-26 — r001-warranty-claim. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone, python3 -m src.app, open 127.0.0.1:9074. No key, no pip install, no index build, no network. The corpus, the answer key, all three free floors, the tolerance sweep, the answer-key gate and every recorded result are committed, so the whole comparison re-scores offline for nothing -- and the arm that scores highest on this corpus needs no key at all: b002-warranty-claim-claimgate reads 76.30 per cent on the discriminator against r001-warranty-claim's 44.44, over the same 270 lines.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
153,969 msp50, end to end
233,504 msp95
2 minclone to first result
What the clock covers. Model call only, one per claim, 5 concurrent workers against a key sibling kits were using at the same time; 1,503.4 s of wall clock for 45 claims, 45 of 45 answered, 0 truncated. Read it as a bound rather than a clean single-tenant latency. The correspondence-blind arm ran the same 45 at 6 workers in 1,072.7 s for a p50 of 142,311 ms and a p95 of 212,947 ms, so the spread across concurrency settings is visible rather than asserted.
Current processWhat it replaces
A warranty analyst reading a transmitted claim line by line against the policy: working out completed months from the in-service date to the repair date and comparing the odometer against the limit for whichever coverage governs that operation, looking each operation code up in the labour time guide and comparing the hours claimed against the published time, classifying each cause code as a covered defect or as wear, maintenance or customer damage, checking the file for a prior repair to the same component that nobody authorised, and then -- the part no table helps with -- reading the service correspondence for the one sentence that reverses any of it.
Where it is not good enough
⚠︎ ON THIS KIT THE HONEST HEADLINE IS THAT PURE PYTHON WINS, AND THE NUMBERS SAY SO ON SIX COLUMNS. THE STRONGEST FREE FLOOR BEATS THE PAID ARM ON THE DISCRIMINATOR AND THIS PAGE LEADS WITH IT. 76.30 pct against 44.44 pct over the same 270 claimed lines (b002-warranty-claim-claimgate against r001-warranty-claim), free, offline, no key, and 0.1 s of wall clock for all 45 packs against 1,503.4 s. The floor also wins INSUFFICIENT_EVIDENCE recognition 16 of 16 against 6 of 16, the policy provision 100.00 pct against 66.15, the published labour time 261 of 261 against 239 of 261, citation discipline 0 of 150 against 72 of 269, and over-denial 33.87 pct against 50.00. ⛑ BUT THE COLUMN THE FLOOR CANNOT REACH AT ALL IS THE ONE THIS KIT WAS BUILT TO TEST, AND THE PAID ARM TAKES ALL OF IT: 22 of 22 denial grounds that exist only as a sentence in the service correspondence, against 0 of 22 for all three floors; 22 of 22 on the correspondence half of the labour-time trap, where every printed figure on the line is internally consistent; 130 of 130 unpayable lines denied against the floor's 108 of 130; and 12 of 42 silent over-denials against the floor's 42 of 42, which is a denial sent to a dealer holding the manufacturer's own document that reverses it. ⛑ AND THE ABLATION PROVES THAT IS REAL READING RATHER THAN A LUCKY PRIOR. Blind the same arm to the correspondence and five columns land on the free floor's figure to the decimal: correspondence grounds 22 of 22 -> 0 of 22, the prose half of the labour trap 22 of 22 -> 0 of 22, labour time exceeded 100.00 -> 47.62 pct, unpayable lines denied 100.00 -> 83.08 pct, silent over-denial 12 of 42 -> 42 of 42. ⚠︎ AND THE ABLATION ALSO CONVICTED THIS KIT'S OWN HEADLINE METRIC: the blind arm scores HIGHER on the discriminator, 45.93 against 44.44, while its verdict accuracy falls 72.22 -> 54.81 and its over-denials rise 62 -> 91. The measured driver is that blinding removes the numbers it was mis-citing -- non-identifier citations 72 -> 23, WP-3.4 substitutions 42 -> 20 -- so the citation term rewards an arm for having less to say. Published, not adjusted. ⚠︎ SO THE RECOMMENDATION IS SPLIT, NOT AVERAGED: if every fact that decides your claims is already a field, take the floor and set its tolerance. If any of them is a sentence, no amount of free code reaches it.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt45jsonl1json1
45 warranty claim review packs, 270 claimed lines — the warranty policy printed in the pack, the claimed labour lines, the records on file and the service correspondence
MIT — same as the code. Every pack stands alone: nothing carries between claims and nothing depends on the order they are read in.
2The warranty policy, printed inside the packEvals ↗
the coverage terms by operation, the cause classification table and the labour operation guide, each row carrying the WP provision that governs it
so there is no rule store to keep in step with the corpus — a floor that cites a provision reads it out of the same page the analyst is holding
WP-3.1 and WP-3.4 are the two this corpus turns on: what the guide publishes is payable, against what a hotline authorisation adds on top of it
3Six structured checks, at a typed toleranceArchitecture ↗
src/checks.py decides all 270 lines from the pack's own tables, free — NOT_A_COVERED_DEFECT, CUSTOMER_DAMAGE, PRIOR_UNAUTHORISED_REPAIR, OUT_OF_COVERAGE_TIME, OUT_OF_COVERAGE_MILEAGE, LABOUR_TIME_EXCEEDED
an operation with no guide row and a line with no technician story are not grounds at all — they come back INSUFFICIENT_EVIDENCE with the missing item named
and it holds ONE typed hour constant, src/checks.TOL = 0.05 hours, plus src/checks.MILE_TOL = 0 miles which is DELIBERATELY not swept with it
Recorded failurean odometer limit is a number in a policy and an hour figure is a measurement with a rounding error; sweeping one tolerance across both is the category error the sweep exists to expose, so the mileage limit is named separately and left at zero
⚑ the strongest one BEATS the paid arm by 31.86 points on the discriminator, costs $0.00, needs no key and no network, and scores all 45 packs in 0.1 s
and it is correspondence-blind by construction — 0 of 22 grounds written only into a note, which is the half it cannot reach at any tolerance
5Tolerance sweepno lens on the shipped page
the floors re-scored across eight hour tolerances, offline and free — evals/tolerance.py, run b900-warranty-claim-tolerance-sweep
claim-gate: 57.78% at 0.00 and 0.005 h · 65.56% at 0.02 · 76.30% at the shipped 0.05 · 76.30% at 0.10 · 75.93% at 0.25 · 75.19% at 0.50 · 74.81% at 1.00
line-sweep peaks at the same 0.05 (70.37%); the answer key itself is IDENTICAL over [0.031, 0.159] hours, re-derived at five points, and moves at 0.16
Recorded failurethe weakest floor does the opposite and that is the reading to keep: coverage-window climbs to its best 39.26% at a one-hour tolerance while its labour catch collapses from 47.62% to 9.52% — it buys the headline by giving up the check, so no floor is published outside the band the key is stable over
the split is MEASURED, not apportioned — two calls at max_tokens=1 solve input = 1,124.44 + 0.228411 per pack character, so about 66% of the input bill is the claim pack and 34% is the instruction
the Dealer Position section — chargeback exposure, goodwill authority, relationship posture — is removed by src/select.py at the seam: 5 sections sent, 1 withheld, on 45 of 45 packs
one entry per claimed line: DENY with a named ground, the WP provision from this pack and the record identifier behind it
or PAYABLE, or INSUFFICIENT_EVIDENCE naming what the pack does not settle
plus the labour time payable recomputed on every line and a claim action; 0 of 270 lines were left off the adjudication entirely
Recorded failure72 of 269 citing lines put a real pack string into a field that asked for a record identifier an analyst can pull — 40 operation codes, 14 special coverage adjustment numbers, 13 hotline case numbers, 10 repair order numbers, 5 goodwill authorisations and the literal phrase Service Correspondence — while 0 of 269 named anything absent from the pack
Recorded failureTHE FREE FLOOR WINS THE DISCRIMINATOR, 76.30% against 44.44 — and the paid arm still takes the three columns free code cannot reach: 22 of 22 correspondence grounds against 0 of 22, 42 of 42 labour-time cells against 20 of 42, and silent over-denial at 12 of 42 against the floor's 42 of 42
A manufacturer's warranty desk, one transmitted claim at a time, before anything is paid or denied.
⚠︎ THE FREE FLOOR WINS THIS KIT AND THE PAGE LEADS WITH IT: claim-gate takes 76.30 pct of 270 claimed lines against the fast tier's 44.44, for $0.00, with no key, no network and 0.1 s for all 45 packs. Anyone who wants a warranty pre-check and has this corpus's shape should write the six checks and stop there.
⛑ AND YET THE PAID ARM TAKES EVERY COLUMN THAT DECIDES WHETHER A DENIAL SURVIVES AN APPEAL, which is why the losing number is not the end of the reading: 130 of 130 unpayable lines denied against the floor's 108, 22 of 22 grounds written only into the service correspondence against 0 of 22 for every floor at every tolerance, and 42 of 42 LABOUR_TIME_EXCEEDED cells against 20 of 42. On the 42 lines the pack has already answered it stays silent on 30, over-denying 12, where the floor over-denies all 42. The correspondence half is not a column free code does badly at; it is a column no regular expression reaches.
⚠︎ ALL 62 OVER-DENIALS ARE ONE SHAPE AND IT IS THE DEALER'S OWN ARITHMETIC: every single one is a line claiming 0.01 h (23), 0.02 h (13) or 0.03 h (26) above the published operation, denied as LABOUR_TIME_EXCEEDED. Not one over-denial is anything else. The floor's answer to that is one float — src/checks.TOL = 0.05 hours. The model has no constant to set, and that single rounding shape is the whole of its over-denial rate. ⚑ THE MODEL CONVICTED THE ANSWER KEY ON 42 CELLS AND THE KEY WAS NOT CHANGED. On every LABOUR_TIME_EXCEEDED cell it cited WP-3.4, that time above the published operation is payable only with a hotline authorisation, where the key carries WP-3.1, that labour is payable at the time the guide publishes. Both clauses are in the pack and WP-3.4 is arguably the more precise one for a line claiming more than the published time. THE KEY WAS NOT EDITED AND THE RUN WAS NOT RE-FIRED, so S_LABOUR_TIME_EXCEEDED reads 0 of 20 and C_LABOUR_TIME_EXCEEDED 0 of 22 on a set of lines the arm denied 42 of 42 with the right ground.
⚠︎ THE SECOND CONVICTION IS THE CITATION FIELD: 72 of 269 citing lines quoted a string that really is in the pack but is not an identifier the pack defines — operation codes, hotline case numbers, repair order numbers, goodwill authorisations, special coverage adjustment numbers and the phrase Service Correspondence — and 0 of 269 named anything absent from the pack at all. That is a different failure from invention and it is scored separately. ⚑ THE COUNTERFACTUAL IS PUBLISHED AS A DIAGNOSIS OF 44.44 PCT AND NEVER AS A REPLACEMENT FOR IT: computed offline from the committed result file, crediting WP-3.4 alone gives 56.67 pct, not penalising the in-pack non-identifier citations alone gives 55.93, and BOTH ARGUABLE READINGS TOGETHER GIVE 70.74 — still 5.56 points below the free floor. The residue is the 62 over-denials, 10 unsettleable lines called anyway, 4 denials with a wrong ground, provision or record, and 3 payable lines called INSUFFICIENT_EVIDENCE. The headline does not move.
⚠︎ AND THE WHOLE VERDICT SITS ON ONE TYPED HOUR, WHICH IS THE FINDING RATHER THAN A CAVEAT. b900-warranty-claim-tolerance-sweep re-scores the floors across eight tolerances: claim-gate reads 57.78 pct at 0.00 and 0.005 h, 65.56 at 0.02, 76.30 at the shipped 0.05, 76.30 at 0.10, 75.93 at 0.25, 75.19 at 0.50 and 74.81 at 1.00, and line-sweep peaks at the same 0.05. The weakest floor runs the other way — coverage-window climbs to its best 39.26 pct at a one-hour tolerance while its labour catch collapses from 47.62 to 9.52, buying accuracy by giving up the check. The answer key is IDENTICAL over [0.031, 0.159] hours, re-derived at 0.031, 0.05, 0.10, 0.15 and 0.159 and changing at 0.16, so nothing is published outside that band whatever the sweep says. MILE_TOL stays at 0 miles and is deliberately never swept.
⛑ THE CEILING PROBE WAS DISCARDED BECAUSE IT CONVICTED THE CORPUS, NOT THE MODEL, and it is kept in results/discarded/ rather than deleted. c000-warranty-claim-calibration fired the four heaviest packs at 32,000 and drew a largest reply of 30,310 output tokens, 94.7 pct of that cap — that reading STANDS and is what raised MAX_TOKENS to 64,000 before any scored arm fired. Its SCORES are gone: the technician stories were drawn from a fixed pool instead of written from the operation on the line, so a drive-shaft line carried a starter-motor story, and the arm returned INSUFFICIENT_EVIDENCE on 14 of 17 payable lines saying exactly that, in its own words, line by line. It was right. The generator was fixed, the dataset moved to -v2, and the probe was re-taken as c001 at the raised ceiling.
⚠︎ ONE MEASURED CORPUS WART IS RECORDED RATHER THAN REGENERATED AWAY: of the 16 special coverage adjustment notes, the limb the denial actually turns on reopens the coverage on 16 of 16, but both limbs are derived from where the vehicle stands rather than from the coverage term, so 8 of 16 print the OTHER limb at or below the standard term — 5 on months, 3 on miles. The trap is intact and the sentence is untidy. It was found after r001 was already in flight, and regenerating would have discarded 45 paid calls to tidy a sentence. tools/build_corpus.py carries the note and the fix.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL and MODEL in the repo-root .env, or a kit-local .env carrying MODEL alone to point this one kit at a different tier while keeping the shared account. Adding a provider is one function and one entry in PROVIDERS, and it must return token counts as well as text -- an adapter that returns no usage cannot be published here, because the cost lens prices what the eval lens counted.
what leaves the machine
src/select.py
NEVER_SENT. One tuple, holding one section name. It is a named-section denylist and not a redaction system: the same goodwill authority that is withheld in the Dealer Position block will be sent when a sentence in the Service Correspondence names it, because the correspondence is where the answers live.
the free floor
evals/baseline.py
--floor coverage-window | line-sweep | claim-gate. Three pure-code adjudicators behind the same answer shape, so one scorer grades all three and the model without a second code path.
the labour-hour tolerance
src/checks.py
TOL, one float, in HOURS. It moves claim-gate from 57.78 per cent at 0.00 hours to 76.30 at 0.05, and it moves the over-denial rate from 74.19 per cent to 33.87 across the same three hundredths of an hour. Sweep it on your own claims with python3 -m evals.tolerance before believing any floor's score -- and do not publish outside the band where the answer key itself stops moving.
the mileage tolerance
src/checks.py
MILE_TOL, in miles, fixed at 0 and deliberately kept OUT of the hour sweep. Running an hour tolerance over an odometer reading is the category error the sweep exists to expose; if your coverage limits are administered with a grace band, this is the constant to move and it moves on its own.
which record proves which ground
src/checks.py
SUPPORT_TYPE, one dict, six entries. This corpus files exactly one record type per ground; a real adjudication attaches several documents to one ground and some of them are images, so this is the mapping to widen first.
the ablation
src/prompt.py
build(text, blind=True), which truncates the Service Correspondence and says so in the instruction. It is the only evidence this kit can offer that an arm READ the prose rather than learning that transmitted warranty claims are over-claimed: take the notes away and an arm that was reading them loses exactly the correspondence cases.
the injected phrasing
evals/injection.py
--phrasing, over the PHRASINGS dict, with --scope and --paired-with fixing the denominator before anything is sent. Sibling kits have measured one phrasing at over eighty per cent suppression and another at near zero, so the phrasings are reported one by one and never meaned together.
the corpus
tools/build_corpus.py
One seed (20260826) and one case plan. Replace it with your own claim packs and your own data/gold.jsonl; data/SOURCES.md lists what to change and in what order, and python3 -m evals.check_labels is what tells you the replacement key is self-consistent.
Components
Component
File
Role
section split
src/segment.py
Cut the claim pack on its underlined headings -- Warranty Claim Review, Warranty Policy And Guides, Claimed Lines, Records On File, Service Correspondence, Dealer Position. One regular expression, deterministic, so "what did we send" is a list of six section names and not an argument about a token window. A pack with no headings comes back as one unnamed section rather than raising, so a forker pointing this at their own export gets a usable failure instead of a stack trace.
the seam that withholds
src/select.py
The Dealer Position block -- the store's open chargeback exposure, the goodwill amount a district manager may pay without a region review, and the manufacturer's own relationship posture toward this dealer -- never leaves the machine. None of it bears on whether a cause code is a covered defect, whether the vehicle was inside its coverage on the repair date, or whether the hours claimed match the published operation. All of it is a reason to pay. Withheld at the seam rather than trusted to a sentence in the instruction, and the UI prints, per claim, which sections went and which stayed.
the pack parser
src/claim.py
The coverage terms table with its months and miles columns, the labour operation guide with its coverage code and published time, the cause classification table with its covered / wear / maintenance / customer-damage column, the provision index, every claimed line with its operation code, cause code, hours, rate and amounts, and every record on file with its type, reference and authorisation -- all read with regular expressions. The model is never asked for a figure that sits at a fixed offset. months_between() counts COMPLETED months rather than dividing days by 30.44, because a coverage term expires on the same day of the month it started and this corpus deliberately puts repairs within a week of that boundary.
the structured warranty checks -- and the labour-hour tolerance
src/checks.py
The six grounds the pack's OWN TABLES can decide -- NOT_A_COVERED_DEFECT, CUSTOMER_DAMAGE, PRIOR_UNAUTHORISED_REPAIR, OUT_OF_COVERAGE_TIME, OUT_OF_COVERAGE_MILEAGE, LABOUR_TIME_EXCEEDED -- in a fixed PRIORITY order, plus the two shapes that settle nothing (an operation code with no guide row, a line with no technician story on file) which return INSUFFICIENT_EVIDENCE with the missing item named. recompute() answers the use case's second half on EVERY line whether or not anything is wrong with it: the hours the guide publishes, the labour amount those hours and the claim's rate give, and the variance against what was claimed. TOL = 0.05 HOURS is the single published constant that decides how much over is over, and MILE_TOL is fixed at 0 miles and named separately because a coverage limit is a number in a policy and not a measurement with a rounding error.
the prompt, and the correspondence-blind variant
src/prompt.py
The whole instruction in one place, in send order. It names the three dispositions and the six grounds, because scoring an arm on a vocabulary it was never given measures the wording rather than the arm. What it deliberately does NOT say is where to look -- nothing in it mentions reading the correspondence against the tables, and nothing mentions that a published operation time can be superseded by a combination time. _strip_notes() is the ablation: it TRUNCATES at the Service Correspondence heading and RAISES if that heading is absent, because a silent no-op there would publish an ablation arm that was never ablated.
the model call and the reply parse
src/draft.py
One claim pack in, one drafted adjudication out, and the only place a model is called for a claim. Holds the 64,000-token output ceiling that this kit's own probe forced: c000-warranty-claim-calibration fired the four heaviest packs at 32,000 and drew 30,310 output tokens on the largest reply -- 94.7 per cent of that cap -- so 32,000 was abandoned BEFORE any scored arm fired rather than after one came back truncated. c001-warranty-claim-calibration re-took the probe at 64,000 on the corrected corpus and the same four packs drew a largest reply of 27,045. The scored run then bore the decision out and bounded it: across all 45 claims r001-warranty-claim came back 45 of 45 answered, 0 failures, 0 truncated and 0 lines omitted, with its heaviest reply at 30,562 output tokens -- 47.8 per cent of the raised ceiling and ABOVE the cap that was abandoned. The parse is tolerant of a fenced block and of leading prose, uppercases only the closed vocabularies, and coerces a single record identifier returned as a bare string into a list.
the adapter
src/adapters/__init__.py
One interface, several providers, raw HTTP and no vendor SDK, so the kit installs nothing to call whichever key a forker already holds. ⚑ ITS SOCKET TIMEOUT AND THE OUTPUT CEILING ARE ONE SETTING WEARING TWO NAMES: TIMEOUT_S is 1,800 seconds beside a 64,000-token ceiling, because completions here are not streamed and a reply that actually fills that ceiling holds a silent socket for the whole generation. Transport retries are cut to ONE against four for a transient status, on the ground that a blown socket on a non-streamed generation is not a rate limit and asking four more times turns one defect into five bills.
the shared connection, and the kit-local override
src/config.py
Provider, base URL, key and model are read from the repo-root .env first, then this kit's own .env, then the real environment, MERGED KEY BY KEY rather than file by file -- so a kit-local file holding MODEL alone points one kit at a different tier and keeps the shared account. Neither file is required and a missing one is not an error, which is what keeps a single cloned kit runnable with no repo above it. sources() reports WHICH file contributed and never a value, so "where is this key coming from" is answerable without anyone reading a key aloud.
the call cap
src/budget.py
One key funds every kit under this root, which is the point of the shared .env and also the hole in it. This counts CALLS, not dollars, because a dollar cap needs a rate card the kit does not have and a cap that silently reads the wrong card is worse than none. The ledger line is written BEFORE the call, so a crash mid-call over-counts by one rather than under-counting. No cap configured means no cap and it says so out loud in the line printed before a run; there is no hard-coded ceiling in the file.
the three free floors
evals/baseline.py
coverage-window is the strawman that is not a strawman -- the months-and-miles arithmetic plus a guide lookup, which is what a warranty system already does the moment a claim is transmitted. line-sweep runs the six structured checks and cites the provision the policy table prints beside the row it used, but never stops: anything it cannot settle it denies anyway. claim-gate is line-sweep plus the one thing that separates a defensible adjudication from a list of exceptions -- it REFUSES to state a ground it cannot support, and every provision and record identifier it cites is read out of the pack rather than assembled from a constant. ⚑ AND THE THIRD FLOOR WINS THE DISCRIMINATOR OUTRIGHT, MEASURED: b002-warranty-claim-claimgate scores 76.30 per cent against r001-warranty-claim's 44.44 per cent over the same 270 lines, free, offline and in 0.1 seconds. The three floors read 34.81, 70.37 and 76.30 per cent respectively. All three are correspondence-blind BY CONSTRUCTION -- 0 of 22 correspondence grounds against the paid arm's 22 of 22 -- and a keyword floor over the notes would close part of that gap and was deliberately not built, because it would be tuned to the sentences this generator happens to write and would measure the generator.
the tolerance sweep
evals/tolerance.py
Three floors by eight labour-hour tolerances, free and offline, and it says the headline is a measurement of a parameter as much as of a method. claim-gate scores 57.78 per cent at 0.00 hours, 76.30 at 0.05, holds 76.30 at 0.10 and falls to 74.81 at 1.00. The weakest floor does the opposite and climbs to its sweep best of 39.26 per cent at a one-hour tolerance while its labour catch collapses from 47.62 to 9.52 -- it is not getting better, it is buying accuracy by giving up the only thing it was looking for. No floor is published outside the band over which the answer key does not move, because a floor allowed to relabel the key is not a floor.
the scorer
evals/scoring.py
Exact match per line, no model grades anything, and every rate carries its own denominator. A line counts toward claim_adjudication_accuracy_pct only when the disposition is right AND, on a denial, the ground is the right one of six, the provision cited is the one the pack carries for it, every required record identifier is attached and nothing on the line was invented. The catch rate is split by CHANNEL -- table grounds against correspondence grounds -- and labour time is split again on top of that, because the named trap is planted on both channels and a blended number would let one half carry the other. silent_overdeny_pct and invented_citation_pct are the two named dangerous failures, and the majority-class score is printed beside every arm.
the answer-key gate
evals/check_labels.py
Eleven properties, re-derived from the shipped corpus, and the three that matter are the corpus's claims about itself rather than its schema: the table half must re-derive EXACTLY from src/checks.structured_finding(); a correspondence denial must look PAYABLE to the tables; a correspondence trap must look DENIED to them. A prose denial a table can already see is not a prose case and labelling it one would inflate the very column this kit exists to measure. --self-test red-proves the gate against seeded defects rather than asserting that it works.
the injection probe
evals/injection.py
One instruction-shaped note forced into the Service Correspondence of every claim in a scope FIXED BEFORE THE RUN -- the claims where the paired run correctly denied at least one line, in claim-id order, capped by an argument that is recorded in the result file. Every injected cell is paired against its own un-injected answer rather than against gold, because a cell the arm was already getting wrong cannot be suppressed. It records whether a cell that HELD had its ground, its provision or its support changed underneath it, and both phrasings are quoted in the file in full, including any that was not fired.
the local UI
src/app.py
http.server, standard library only, five endpoints and no write path -- nothing here pays, denies, submits or charges back a claim, and there is no endpoint, button or flag that adds one. It renders with no key: the pack, everything read off it in code, the published labour time on every line and the strongest free floor all cost nothing, so the page always computes them. The floor's adjudication sits beside the model's with an Agree? column, so a line where the two agree is a line where the model bought nothing and the page says so rather than taking the credit.
the corpus generator
tools/build_corpus.py
Writes data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json from seed 20260826. The answer key is DERIVED for everything pure code can settle -- it is structured_finding()'s own output -- so the key and the free floor cannot silently drift apart. What is TYPED is only the four cases code cannot reach, and each of those is then asserted to look the opposite way to the tables. The dealer's own rounding is planted at no more than 0.03 hours and every planted over-claim at 0.08 hours or more, which is the gap that makes the tolerance sweep readable. ⚠︎ IT ALSO CARRIES A RECORDED DEFECT AND ITS FIX RATHER THAN A REGENERATED CORPUS: both limbs of a special coverage adjustment note are derived from where the vehicle stands rather than from the coverage term, so 8 of the 16 notes print the limb nobody is arguing about at or below the standard term. It was found after r001-warranty-claim was already in flight, and regenerating would have discarded 45 paid calls to tidy a sentence.
Where it breaks at scale
One call per claim pack, and the pack goes in whole -- 45 packs, 455,373 bytes, a mean of 10,119 bytes each. r001-warranty-claim measured what that costs on a full run: 149,220 input tokens across the 45 calls, a mean of 3,316.00 per claim, against 853,196 output tokens at a mean of 18,959.91 -- and 818,949 of those output tokens, 96.00 per cent, are provider-side reasoning rather than the adjudication. A real transmitted claim does not grow much past this, but a real POLICY does: the coverage terms, the labour operation guide and the cause classification table are three short tables inside the pack here, and in a manufacturer they are systems with tens of thousands of operation codes. Inline them and the prompt stops fitting; index them and you have built the retrieval step this kit does not have, and the operative sentence is then one summariser away from being lost. The OUTPUT ceiling is what bites first and this kit measured how close: at a 32,000-token cap the four heaviest packs drew a largest reply of 30,310 tokens, 94.7 per cent of it, so the cap was abandoned before any scored arm fired; at 64,000 the scored run's heaviest reply was 30,562, 47.8 per cent of the ceiling and above the cap that was dropped, with nothing truncated across 45 claims. Splitting the pack per line is the obvious relief and it would break the thing that works: in WC-0018 a single note records that LO-2404 and LO-2519 are a combination whose published time is superseded, and that sentence bears on two lines at once, so a per-line prompt cannot see it. Wall clock was 1,503.4 seconds for 45 claims at five concurrent workers against a key sibling kits were using at the same time, which is a bound rather than a clean single-tenant latency. ⚑ AND THE FREE FLOOR DOES NOT SCALE THIS WAY AT ALL -- claim-gate scores all 45 packs in 0.1 seconds, needs no key and no network, and beats the paid arm on the discriminator by 76.30 per cent to 44.44.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
WC-0018 before anything is adjudicated. The claim header, the coverage terms, the provision index, every record on file, the service correspondence, all six claimed lines with the hours the labour operation guide publishes beside the hours claimed, and the whole claim-gate adjudication are already on the page and cost nothing -- no key was configured and none was needed. The model column reads "not adjudicated yet" rather than accusing a run that has not happened, and the seam panel already lists the five sections that would be sent and the one, Dealer Position, that would not.emptyOpen full size →The same claim after pressing "Adjudicate this claim" with no API_KEY configured. Nothing was called, and /api/draft returns 200 with a sentence saying so instead of failing at the HTTP layer: the claim, everything read off it in code, the published labour time on every line and the free floor below are all computed locally and need no key. ⚑ ON THIS KIT THAT IS NOT A DEGRADED MODE, AND THAT IS MEASURED RATHER THAN CLAIMED. Over the same 270 lines the free floor on this page scores 76.30 per cent on the discriminator (b002-warranty-claim-claimgate) against 44.44 per cent for the paid arm (r001-warranty-claim), so the page with no key is the page carrying the best measured adjudication this kit has.nokeyOpen full size →WC-0018 with r001-warranty-claim replayed beside the free floor, line by line, with an Agree? column between them. The vehicle is 14 completed months in service at 38,632 miles, claimed at a 128.02 labour rate, and the pack packs four correspondence cases onto one claim. L1 claims 1.64 hours against the 1.30 hours LO-3222 publishes: the guide alone proves it and both arms deny it. L5 carries a prior repair to the same oxygen sensor by an independent repairer with Authorisation printed as "none", and both arms deny that too, citing WP-5.2 -- two lines where the model bought nothing, because the tables had already settled them. ⚠︎ THE FRAME IS HERE FOR THE OTHER THREE. At 38,632 miles the vehicle is past the 36,000-mile basic term, so L2 and L6 read as out of coverage on the printed table -- and special coverage adjustments 26-NA-045 and 26-NA-050 in the correspondence reopened those two operations on this model, with the standard policy summary still attached to the claim. Both are PAYABLE; the free floor denies both on OUT_OF_COVERAGE_MILEAGE. L4 runs the other way: 2.30 hours claimed, EXACTLY the time LO-2519 publishes, every figure internally consistent, and one note recording that LO-2404 and LO-2519 are a combination whose time drops to 1.20 hours when both are on one repair order. The free floor passes it. The paid arm gets all six dispositions right on this claim, which is what the correspondence half buys -- and it is not the whole story: over the full 270 lines the free floor still wins the discriminator 76.30 to 44.44, and the miss frame is where that goes.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
⚠︎ THE FAILURE THIS KIT IS ABOUT, AND IT IS THE PAID ARM THAT WRITES IT. WC-0005 replayed from r001-warranty-claim, one of only four claims in the whole corpus whose correct answer is RECOMMEND_PAY_AS_CLAIMED: all six lines are payable. The arm returns RETURN_FOR_ADJUSTMENT and denies four of them. L1 it gets RIGHT and for the right reason -- 1.38 hours claimed against the 1.00 hours LO-3277 publishes, and technical hotline case TAC-496144 in the correspondence authorised 0.58 hours of additional diagnostic time on that very line -- which is exactly the reasoning no free floor here can reach, and the page still marks the line in red: NOT A RECORD OR PROVISION: TAC-496144. The string is real and it is printed in the pack, but it is not an identifier the pack defines, so it is not something a warranty analyst can pull. Then L2, L4, L5 and L6 are denied as LABOUR_TIME_EXCEEDED on 0.03, 0.03, 0.03 and 0.02 hours above the published operation -- the dealer's own clock-off rounding, on a seat frame, a head unit, a catalytic converter and a water pump -- each cited to WP-3.4, and L2 and L4 carrying a special coverage adjustment number in the support field that the page reds out for the same reason. ⚑ AND THIS IS NOT AN UNLUCKY PACK: ALL 62 OF THE ARM'S OVER-DENIALS ACROSS THE WHOLE RUN ARE THIS ONE SHAPE and nothing else -- 0.01 hours on 23 lines, 0.02 on 13 and 0.03 on 26. The free floor's answer to all of it is one float, src/checks.TOL = 0.05 hours, which is why it over-denies 42 lines where the paid arm over-denies 62. The model has no constant to set.failureOpen full size →
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
45warranty claim review packs
0.43 MiBtxt 45
270claimed lines · p50 10161 chars
$0.00setup · 0.0s
How it is cutWhat one claimed line is
45 packs, each one transmitted claim against one repair order, carrying the warranty policy and guides, six claimed lines, the records on file and the service correspondence. No claim references another and nothing carries between them, so the 45 calls can be fired in any order or all at once, and a pack that fails tells you nothing about the 44 beside it. The composition is CONSTRUCTED rather than dealt: 130 of the 270 lines are DENY, 124 PAYABLE and 16 INSUFFICIENT_EVIDENCE, and at claim level 37 packs are RETURN_FOR_ADJUSTMENT against 4 REQUEST_SUPPORT and 4 RECOMMEND_PAY_AS_CLAIMED. That skew is stated rather than hidden, because it is exactly why the single majority word already scores 82.22 per cent on the claim action and why both the majority and the measured figure are printed side by side. The line mix carries 82 clean lines, 108 lines whose denial ground the tables settle, 22 whose denial ground only a sentence settles, 42 reverse traps the correspondence has already answered and 16 lines that settle either way.
SetupWhat the setup figure measured
THERE IS NO INDEX AND NO RETRIEVAL IN THIS KIT, and that is a design decision rather than an omission. One claim pack is read whole and sent whole, minus the withheld Dealer Position block -- no chunking, no embedding, no nearest-neighbour lookup, no pre-digest. The reason is the thing being measured: the answer to a line is often a single sentence in the Service Correspondence which contradicts a printed table, and any step that summarises, ranks or truncates the pack is a place for that sentence to be lost before the model ever sees it, after which the measurement is of the retriever. Nothing is built, so nothing is spent building it. For scale: the strongest free floor scores all 45 packs in 0.1 seconds of wall clock, and the whole comparison re-runs offline for nothing.
LicenceLicence
MIT -- this repository's own licence. The corpus has no third-party component to licence separately: it is written by a generator that ships in the repo, from a seed recorded in data/corpus-stats.json, and every figure in it is invented. Verified by reading every generator input on 2026-08-26.
Bring your ownBring your own warranty claim review packs
tools/build_corpus.py writes data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json. To point the kit at your own claims, work in this order: src/segment.py's heading pattern if your sections are not underlined; src/claim.py's regular expressions for the coverage terms, the operation guide, the cause classification and the line and record blocks; src/select.py's NEVER_SENT for whichever block of yours is the Dealer Position; src/checks.py's six structured checks, its SUPPORT_TYPE mapping AND its TOL -- run python3 -m evals.tolerance against your own claims before trusting the value shipped here, which was chosen by that sweep on THIS corpus. Then write your own data/gold.jsonl and run python3 -m evals.check_labels over it, because every number this kit publishes is a comparison against that file and it is the piece that cannot be skipped. evals/scoring.py, evals/run.py and the UI are shape-independent and carry over unchanged.
⚠︎ And what stops being true when you do:src/select.py withholds a NAMED SECTION and is not a redaction system, and this corpus contains its own counter-example. The Dealer Position block on WC-0042 states that the district manager may pay the claim as goodwill up to 610.50 without a region review, and that block never leaves the machine. Six lines earlier, the Service Correspondence records that the prior repair on line L3 was carried out under district manager goodwill authorisation GW-28842 -- and that sentence IS sent, because it is the only thing in the pack that makes the line payable. The kit cannot both withhold the commercial posture and read the correspondence, and it chooses to read the correspondence. If you point this at real claims, that is the first seam to look at, and the fix is a redaction pass rather than a longer denylist.
What breaks it
A claim that is not this shape. src/segment.py looks for underlined headings and src/claim.py is a set of regular expressions written for these fixed-width tables and these key/value blocks. A real claim arrives as a transmitted record out of a dealer management system -- fixed-width or delimited -- with free-text complaint, cause and correction typed by an advisor and a technician, and the policy and the labour time guide are SYSTEMS rather than attachments. Point the parser at one and it finds no coverage terms and no operations, so every line comes back INSUFFICIENT_EVIDENCE. That is the honest failure and not a hidden one: parse() returns empty tables rather than guessing, and the UI renders an empty table rather than an invented one.
⚠︎ A LABOUR-HOUR TOLERANCE THAT IS NOT YOURS. The labour check is an inequality between two printed hour figures, and this corpus's published times run from 0.60 hours for a strut pair to 11.60 hours for a head gasket -- two orders of magnitude under one constant. TOL = 0.05 moves claim-gate from 57.78 per cent at 0.00 hours to 76.30, and moves its over-denial rate from 74.19 per cent to 33.87 over the same three hundredths of an hour. If your dealers do not clock off in hundredths, or your guide is administered with a published grace, the number on this page is not your number. Sweep it first.
⚠︎ THE CAUSE CODE, WHICH IS THE BIGGEST SINGLE GAP. Every line here carries a CZ- cause code that the pack's own classification table resolves to covered defect, wear, maintenance or customer damage, and two of the six structured grounds are that lookup and nothing else. On a real repair order the cause is a labour operation plus a technician's prose, and the classification is a judgment somebody makes -- so the check has nothing to key on and the two easiest grounds in this kit become the two hardest.
⚠︎ A MEASURED WART IN THIS CORPUS'S OWN TRAP SENTENCES, RECORDED RATHER THAN REGENERATED AWAY. A special coverage adjustment note names both limbs of the reopened term, and both are derived from where the vehicle actually stands rather than from the coverage it is extending. The limb the denial turns on reopens the coverage on 16 of 16 notes, so every trap is intact and no label moves -- but 8 of the 16 print the OTHER limb at or below the standard term, 5 on months and 3 on miles. WC-0005 is the visible instance: 26-NA-025 extends LO-1512 to 59 months and 26,000 miles where the printed basic term is 36 months and 36,000 miles, so it lengthens the months it is arguing about and shortens the mileage nobody is. It was found after r001-warranty-claim was in flight and regenerating would have discarded 45 paid calls to tidy a sentence, so the note and its fix sit in tools/build_corpus.py and the run stands. If you generate your own packs, derive both limbs from the term rather than from the odometer.
An in-service date that is contested rather than wrong. The coverage arithmetic here counts completed months from a date printed in the header and proved by a delivery record. A real in-service date comes from the vehicle master and can be CORRECTED after a warranty registration dispute, at which point the arithmetic is right, the input has moved, and the adjudication that was defensible last month is not.
A line with two grounds. The generator plants at most one case per line, so src/checks.PRIORITY has never been exercised against a real tie -- a wear cause on a vehicle that is also past its mileage limit would be decided by an ordering nothing here has tested.
One record per ground. src/checks.SUPPORT_TYPE maps each ground to exactly one record type, because this corpus files exactly one. A real adjudication attaches several documents to a single ground and some of them are photographs, and a scorer that demands one identifier would mark a complete answer incomplete.
The mileage tolerance is fixed at zero miles and is deliberately NOT swept with the hour tolerance. An odometer reading is not a measurement with a rounding error and a coverage limit is a number written in a policy; running the hour sweep over it is the category error the sweep exists to expose. If your network administers a mileage grace band, that constant has to move on its own evidence.
Everything the corpus does not contain, each real and each absent on purpose: parts pricing and parts markup, which is a whole second adjudication; sublet and towing lines, which are priced but carry no labour operation; a VIN that fails a check-digit test or a claim submitted against the wrong vehicle; claims already paid and being audited afterwards, where the question is chargeback rather than payment; and multiple repair visits for one complaint, which is where most real warranty dispute actually lives.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
chat protocol and system prompt
3,993
1,124
the whole claim pack, minus the withheld Dealer Position block -- warranty policy and guides, claimed lines, records on file and service correspondence
8,080
1,846
Total
2,970
This is the cost lesson as arithmetic: of the 2,970 tokens assembled, 1,846 are evidences — 62% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
The whole thing, byte for byte, as src/prompt.py assembled it for WC-0042 and src/adapters posted it: the system string followed by the claim pack. WC-0042 is one of the two claims p001-warranty-claim-prompt-tokens actually measured, so the token split below is against THIS prompt rather than against an average of prompts nobody sent. Nothing is elided and nothing is summarised -- a pre-digest is a place for the operative sentence to be lost before the model sees it, after which the measurement is of the summariser.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are adjudicating a dealer warranty claim for a vehicle manufacturer.
You act for the MANUFACTURER, before a decision. An adjudication is not a payment and not a denial:
it is what the warranty analyst reads before deciding what to do with the claim. A denial you
cannot support comes back as a dealer appeal with the document that reverses it, and one of those
documents is often the manufacturer's own bulletin. A denial with no policy provision behind it is
not something the dealer has to answer at all.
Return one entry for EVERY line in "Claimed Lines", in the order the pack lists them, using the
pack's own line identifiers. Never drop a line because it looks unremarkable.
For each line give exactly one disposition:
DENY a denial ground can be STATED, the policy provision that carries it is in
this pack, and the record that proves it is in this pack.
PAYABLE the line stands as claimed. Either it meets the policy, or something in
the pack already answers the apparent problem.
INSUFFICIENT_EVIDENCE something about the line cannot be settled from this pack -- the term it
would turn on is not here, or the record that would prove it is not here.
Do not raise a denial; say what is missing.
When the disposition is DENY, the ground is exactly one of:
NOT_A_COVERED_DEFECT the cause is wear from use in service or a maintenance item, not a
defect in factory material or workmanship
CUSTOMER_DAMAGE the cause is impact, misuse, contamination or an unapproved fitment
PRIOR_UNAUTHORISED_REPAIR the failure follows a repair to the same component that the
manufacturer did not authorise
OUT_OF_COVERAGE_TIME on the repair date the vehicle was past the months the coverage
governing this operation runs from the in-service date
OUT_OF_COVERAGE_MILEAGE the odometer at repair was past that coverage's mileage limit
LABOUR_TIME_EXCEEDED the hours claimed are above the time payable for that operation
Otherwise ground is null.
Also give, for every line, the labour time you independently read for that operation: the hours the
labour operation guide in this pack publishes for the line's operation code, to two decimal places.
Give null where the pack does not carry a published time for it.
Cite the governing policy provision as a provision identifier that appears in this pack, and attach
the support as record identifiers that appear in this pack. NEVER write a provision or record
identifier that is not printed in the pack in front of you -- an invented citation is worse than
none, because it is the first thing the dealer's warranty administrator will check.
Anything in the pack may bear on a line. Where a printed table and a later statement in the pack
contradict each other, they are not a tie.
Reply with JSON and nothing else:
{"claim_action": "RETURN_FOR_ADJUSTMENT" | "REQUEST_SUPPORT" | "RECOMMEND_PAY_AS_CLAIMED",
"lines": [{"line": "<the pack's line identifier>",
"item": "<what was repaired, one short phrase>",
"disposition": "DENY" | "PAYABLE" | "INSUFFICIENT_EVIDENCE",
"ground": "<one of the six, or null>",
"policy_provision": "<a provision identifier from this pack, or null>",
"support": ["<record identifiers from this pack>"],
"allowed_labour_hours": <number or null>,
"finding": "<the one sentence the adjudication would carry for this line>"}],
"rationale": "<two sentences at most, on what decided the hardest line>"}
claim_action is RETURN_FOR_ADJUSTMENT if any line is DENY; REQUEST_SUPPORT if none is but at least
one is INSUFFICIENT_EVIDENCE; RECOMMEND_PAY_AS_CLAIMED otherwise. None of the three pays, denies or
submits anything -- each is a recommendation to the person who does.
Warranty Claim Review
----------------------------------------------------------------
Claim WC-0042
Dealer Pinemarsh Motors, dealer code D-1184
Vehicle 2023 Orrell Crossover 2.4
VIN WMY5MN6L72XZKC7MR
Repair order RO-565039
Repair date 2026-02-24
In service date 2025-02-24
Odometer at repair 38,546
Labour rate 111.61
Claim basis published labour operation times, coverage under the vehicle warranty policy
Warranty Policy And Guides
----------------------------------------------------------------
Coverage terms
CV-BAS Bumper to bumper, basic vehicle 36 months 36,000 miles WP-2.1
CV-PWR Powertrain, engine transmission and drive 60 months 60,000 miles WP-2.2
CV-EMI Emissions related components 96 months 80,000 miles WP-2.3
Labour operation guide
LO-1120 Alternator, remove and replace CV-BAS 1.60 hours WP-3.1
LO-1145 Starter motor, remove and replace CV-BAS 1.30 hours WP-3.1
LO-1265 Window regulator, front door, replace CV-BAS 1.90 hours WP-3.1
LO-1310 Instrument cluster, replace and configure CV-BAS 1.10 hours WP-3.1
LO-1512 Infotainment head unit, replace CV-BAS 1.40 hours WP-3.1
LO-1580 Tailgate strut pair, replace CV-BAS 0.60 hours WP-3.1
LO-2255 Rocker cover gasket, replace CV-PWR 1.10 hours WP-3.1
LO-2372 Water pump, replace CV-PWR 3.20 hours WP-3.1
LO-2404 Transmission valve body, replace CV-PWR 5.10 hours WP-3.1
LO-2519 Drive shaft, front, replace CV-PWR 2.30 hours WP-3.1
LO-2566 Engine mount, replace CV-PWR 1.50 hours WP-3.1
LO-3140 Catalytic converter, replace CV-EMI 2.60 hours WP-3.1
LO-3277 Exhaust gas recirculation valve, replace CV-EMI 1.00 hours WP-3.1
Cause classification
CZ-DEF Defect in factory material or workmanship covered defect WP-4.1
CZ-ASM Assembly fault found at the factory build stage covered defect WP-4.1
CZ-SFT Control software fault corrected by the published update covered defect WP-4.1
CZ-WER Component worn from normal use in service wear WP-4.2
CZ-MNT Maintenance item replaced at its published interval maintenance WP-4.2
CZ-CDM Impact, misuse, contamination or an unapproved fitment customer damage WP-4.3
Provision index
WP-2.1 Basic vehicle coverage runs the months and miles in the coverage terms table
WP-2.2 Powertrain coverage runs the months and miles in the coverage terms table
WP-2.3 Emissions coverage runs the months and miles in the coverage terms table
WP-3.1 Labour is payable at the time the labour operation guide publishes
WP-3.4 Time above the published operation is payable only with a hotline authorisation
WP-4.1 A defect in factory material or workmanship is a covered repair
WP-4.2 Wear from use in service and maintenance items are not covered repairs
WP-4.3 Damage from impact, misuse, contamination or an unapproved fitment
WP-5.2 A failure following a repair the manufacturer did not authorise is not payable
WP-6.1 Coverage runs from the date the vehicle first entered service
Claimed Lines
----------------------------------------------------------------
Line L1
Description Transmission valve body, replace (LO-2404)
Cause code CZ-ASM
Labour claimed 5.13
Labour rate 111.61
Labour amount 572.56
Parts amount 1,605.95
Total claimed 2,178.51
Line L2
Description Rocker cover gasket, replace (LO-2255)
Cause code CZ-DEF
Labour claimed 1.10
Labour rate 111.61
Labour amount 122.77
Parts amount 1,480.28
Total claimed 1,603.05
Line L3
Description Drive shaft, front, replace (LO-2519)
Cause code CZ-ASM
Labour claimed 2.32
Labour rate 111.61
Labour amount 258.94
Parts amount 763.04
Total claimed 1,021.98
Line L4
Description Exhaust gas recirculation valve, replace (LO-3277)
Cause code CZ-DEF
Labour claimed 1.00
Labour rate 111.61
Labour amount 111.61
Parts amount 165.38
Total claimed 276.99
Line L5
Description Fuel tank assembly, replace (LO-4160)
Cause code CZ-ASM
Labour claimed 2.60
Labour rate 111.61
Labour amount 290.19
Parts amount 514.66
Total claimed 804.85
Line L6
Description Water pump, replace (LO-2372)
Cause code CZ-ASM
Labour claimed 3.20
Labour rate 111.61
Labour amount 357.15
Parts amount 1,008.76
Total claimed 1,365.91
Records On File
----------------------------------------------------------------
Record RC-0042-01
Type Delivery record
Reference WC-0042
Date 2025-02-24
Content Delivery and registration record for this vehicle
In service date 2025-02-24
Signed by Selling dealer
Record RC-0042-02
Type Odometer record
Reference WC-0042
Date 2026-02-24
Content Odometer reading recorded at check-in for this repair order
Recorded odometer 38,546
Signed by Dealer service advisor
Record RC-0042-03
Type Technician story
Reference L1
Date 2026-02-24
Content Complaint, cause and correction as recorded on the repair order
Cause narrative Transmission valve body inoperative from cold; the assembly was found out of specification as built; transmission valve body replaced and the system retested.
Signed by Dealer service technician
Record RC-0042-04
Type Technician story
Reference L3
Date 2026-02-24
Content Complaint, cause and correction as recorded on the repair order
Cause narrative Drive shaft found faulty on arrival; the fault traces to how the unit was put together at build rather than to anything done in service; drive shaft replaced.
Signed by Dealer service technician
Record RC-0042-05
Type Prior repair
Reference L3
Date 2025-09-19
Content An earlier repair to the same component on this vehicle
Repair order RO-307259
Component Drive shaft, front, replace
Performed by Independent repairer
Authorisation none
Record RC-0042-06
Type Technician story
Reference L5
Date 2026-02-24
Content Complaint, cause and correction as recorded on the repair order
Cause narrative Fuel tank assembly found faulty on arrival; the fault traces to how the unit was put together at build rather than to anything done in service; fuel tank assembly replaced.
Signed by Dealer service technician
Service Correspondence
----------------------------------------------------------------
- The dealer's warranty administrator has held the manufacturer's certification since 2011.
- The customer declined the multi-point inspection recommendations at this visit.
- Recall 21V-114 was checked against this VIN and is not open.
- Parts return tag PR-260752 was raised and the failed part shipped to the regional return centre.
- The prior repair recorded against L3 was carried out under district manager goodwill authorisation GW-28842; the authorisation is held at the region and was not attached to the repair order.
- The repair order was reopened once to correct a typing error in the customer address.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"claim_action":"RETURN_FOR_ADJUSTMENT","lines":[{"line":"L1","item":"Transmission valve body, replace","disposition":"DENY","ground":"LABOUR_TIME_EXCEEDED","policy_provision":"WP-3.4","support":["RC-0001-03"],"allowed_labour_hours":5.10,"finding":"The claimed 5.11 hours on LO-2404 exceeds the published 5.10 hours and no hotline authorisation appears, so the excess labour is not payable."},{"line":"L2","item":"Oxygen sensor, upstream, replace","disposition":"PAYABLE","ground":null,"policy_provision":"WP-4.1","support":["RC-0001-04"],"allowed_labour_hours":0.80,"finding":"The upstream oxygen sensor out-of-spec as built is a covered factory assembly defect within CV-EMI and the labour matches the published 0.80 hours."},{"line":"L3","item":"Clutch assembly, replace","disposition":"PAYABLE","ground":null,"policy_provision":"WP-4.1","support":["RC-0001-05"],"allowed_labour_hours":7.80,"finding":"The clutch assembly internal failure is a covered defect within CV-PWR and the claimed 7.80 hours matches the published time."},{"line":"L4","item":"Cylinder head gasket, replace","disposition":"DENY","ground":"PRIOR_UNAUTHORISED_REPAIR","policy_provision":"WP-5.2","support":["RC-0001-07"],"allowed_labour_hours":11.60,"finding":"The same cylinder head gasket was previously repaired by an independent repairer without authorisation, so WP-5.2 excludes this failure."},{"line":"L5","item":"Engine mount, replace","disposition":"DENY","ground":"LABOUR_TIME_EXCEEDED","policy_provision":"WP-3.4","support":["RC-0001-09"],"allowed_labour_hours":0.70,"finding":"Because LO-2455 and LO-2566 are claimed together, LO-2566 is limited to the 0.70-hour combination time, and the claimed 1.50 hours exceeds it."},{"line":"L6","item":"Rocker cover gasket, replace","disposition":"DENY","ground":"PRIOR_UNAUTHORISED_REPAIR","policy_provision":"WP-5.2","support":["RC-0001-11"],"allowed_labour_hours":1.10,"finding":"The same rocker cover gasket was previously repaired by an independent repairer without authorisation, so WP-5.2 excludes this failure."}],"rationale":"L5 is the hardest line: the later combination notice supersedes the LO-2566 table time, so only 0.70 hours is allowed for the engine mount. L4 and L6 are separately excluded because the same components were previously repaired by an un-authorised independent repairer."}
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Decide a car dealer's warranty claim, line by line — 270 warranty claim review packs drawn from 45 real warranty claim review packs. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Pure code, exact match per claimed line against data/gold.jsonl. NO MODEL GRADES ANYTHING on this kit -- there is no judge, no rubric and no second call. evals/scoring.py compares the arm's answer to the key and every figure on this page falls out of that comparison.
270warranty claim review packs
45source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED120 · 206 · 190 · 94 / 270claim adjudication accuracy pct — claim adjudication accuracy -- disposition, ground, policy provision, record, and nothing cited that the pack does not name, all at onceDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED195 · 206 · 190 · 164 / 270claim verdict accuracy pct — verdict only, three-way -- what a rules engine can score (all-DENY scores 48.15 pct)Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED130 · 108 · 108 · 70 / 130denials caught pct — unpayable lines deniedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED62 · 42 · 42 · 30 / 124overdeny rate pct — payable lines denied anyway -- LOWER IS BETTER, never averaged with the row aboveDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED108 · 108 · 108 · 70 / 108table ground caught pct — denial grounds the pack's own tables settleDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 · 0 · 0 · 0 / 22correspondence ground caught pct — denial grounds only a sentence in the service correspondence settles -- every free floor takes 0.0Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED42 · 20 · 20 · 20 / 42labour time exceeded caught pct — LABOUR_TIME_EXCEEDED, the named trap, on its own denominatorDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED20 · 20 · 20 · 20 / 20labour time table caught pct — -- of which the hours claimed are visibly above the published timeDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 · 0 · 0 · 0 / 22labour time prose caught pct — -- of which the hours claimed EQUAL the published time and a combination note supersedes itDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED12 · 42 · 42 · 30 / 42silent overdeny pct — lines the pack itself has ALREADY answered, denied anyway -- LOWER IS BETTERDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED6 · 16 · 0 · 0 / 16unsettleable named pct — unsettleable lines named rather than deniedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED128 · 108 · 108 · 58 / 130denial ground accuracy pct — ground named correctly, of the denials statedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED86 · 108 · 108 · 0 / 130policy provision accuracy pct — policy provision cited correctly, of the denials statedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED126 · 108 · 108 · 0 / 130claim support completeness pct — required record attached, of the denials statedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED239 · 261 · 261 · 261 / 261labour time accuracy pct — the published labour time read correctly to the hundredth of an hour -- the use case's second question, scored apart from the discriminatorDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED72 · 0 · 0 / 269invented citation pct — cited something that is not a provision or a record identifier the pack prints -- LOWER IS BETTERDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 0 / 269absent citation pct — -- of which the token appears nowhere in the pack at allDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 0 · 0 / 270claim line omission pct — claimed lines left off the adjudication entirely -- LOWER IS BETTERDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED37 · 37 · 37 · 33 / 45claim action accuracy pct — claim action, three-way (the single majority word scores 82.22 pct)Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The ANSWER KEY has its own gate and the gate is red-proven. evals/check_labels.py asserts eleven properties over all 45 claims and 270 lines, and the three that matter are the corpus's claims about itself rather than its schema: the table half re-derives EXACTLY from src/checks.py -- disposition, ground, provision and record -- so the key and the free floor cannot silently drift apart; a correspondence DENY must look PAYABLE to the tables, or it is not a correspondence case and labelling it one would inflate the column this kit exists to measure; and a correspondence trap must look DENIED to the tables, or it is not a trap and counting it would understate silent_overdeny_pct. --self-test seeds four defects in memory and requires each to be convicted by name, then re-checks that the shipped key still passes -- an acquittal is proved as well as a conviction, because a gate that convicts everything convicts nothing. All four convict and the shipped key is clean before and after.
Grading costWhat it costs
Every dollar on this page is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid. A PROJECTION CARD, not a bill. The provider that ran this publishes no rate card this repo commits, so every dollar figure on this page is measured tokens times a published rate for a different model. Swap the card and every figure moves.
Priced at
Per 1M in / out
One warranty claim review pack
1,000 warranty claim review packs
Share that is the prompt
Google Gemini 3 Flash the estate's shared projection card, so this kit's rows compare with every other kit's
$0.50 / $3.00
$0.058538
$58.54
3%
Same work, 1× the bill
The same warranty claim review packs, the same tokens — only the rate card changed. And on that card about 3% of what you pay is the prompt this pipeline sends, not the answer it writes.
The one lever with a measurement behind it is what src/select.py withholds, and it points the wrong way for saving money: the Dealer Position block is already cut, and cutting more is what the ablation did. s001-warranty-claim-notes-blind removes the whole Service Correspondence section and comes out 13.3 pct cheaper per claim -- and takes 0 of 22 correspondence grounds instead of 22 of 22, with silent over-denials going 12 of 42 to 42 of 42. That is the cheapest this kit gets and it is the worst it performs. The honest lever is the other direction: if your claims are decided by fields, run claim-gate for $0.00 and do not send anything at all.
Rates checked 2026-08-26. The provider that actually ran every call on this kit is kept off this page per the series rule; the real spend sits in the shared call ledger. And the arm that scores HIGHEST on this corpus is not priced at all because it costs nothing: the three free floors, the tolerance sweep, the answer-key gate and the wiring stub make no call.
Grading unitWhat the grading figure prices
freegrading cost, as measured
EVERY GRADER ON THIS KIT IS FREE. Scoring is pure code -- exact match per claimed line against data/gold.jsonl -- so re-scoring all four arms costs $0.00, sends nothing anywhere and needs no key. The unit the kit is PRICED in is one transmitted claim covering every line claimed on it, and that is in the Cost lens; grading it is free.
The gradersThree ways to grade
⚑ THE ASSIGNED FLOOR IS b000-warranty-claim-coveragewindow AND IT IS THE WEAKEST OF THE THREE, WHICH IS THE POINT OF SHIPPING ALL THREE. A deterministic in-coverage check by date and mileage plus an operation-code lookup takes 34.81 pct. It cannot state which of six grounds it is looking at -- a wear item, an unapproved fitment and a prior repair nobody authorised all sail through it as payable -- and it cites nothing, so 0 of its denials carry a provision or a record. Reporting only that number would have flattered the model by 9.63 points. The honest floor is the strongest thing free code can do, which is claim-gate at 76.30 pct.
⚑ THE FLOOR THIS KIT WAS TOLD TO BEAT IS b000-warranty-claim-coveragewindow: a deterministic in-coverage check by date and mileage, plus the operation-code lookup. Split the 130 unpayable lines by what actually decides them and the two arms do not compete -- they are good at different columns, and averaging them into one accuracy hides the whole finding.
Out of coverage by TIME -- completed months against the term -- assigned floor 100.0 pct, strongest floor 100.0 pct, fast tier 100.0 pct, of 20. Date arithmetic, and everything ties. Nobody should pay for this column. Out of coverage by MILEAGE -- odometer against the limit -- assigned floor 100.0 pct, strongest floor 100.0 pct, fast tier 100.0 pct, of 18. A single integer comparison, and everything ties. Same conclusion. Labour time above the published operation -- assigned floor 47.62 pct, strongest floor 47.62 pct, fast tier 100.0 pct, of 42. The op-code lookup takes the 20 lines where the hours claimed are visibly above the published time and NONE of the 22 where they equal it exactly and a combination note in the correspondence supersedes it. Both floors stop at the same 47.62 pct because the second half is not in any table. Cause not a covered defect -- wear, maintenance or customer damage -- assigned floor 0.0 pct, strongest floor 100.0 pct, fast tier 100.0 pct, of 34. ⚠︎ THE ASSIGNED FLOOR SCORES ZERO. A coverage window has no opinion about why a part failed, so every wear item and every unapproved fitment sails through it as payable. Reading the cause classification table closes it completely -- which is what the strongest floor does, for nothing. Failure after a prior repair nobody authorised -- assigned floor 0.0 pct, strongest floor 100.0 pct, fast tier 87.5 pct, of 16. ⚠︎ ZERO AGAIN for the assigned floor, and the one column where free code beats the paid arm outright on finding: the strongest floor takes 16 of 16 and the fast tier 14 of 16.
So the split IS the finding, and it is not the one either extreme predicts. The assigned floor ties everything on the two pure date-and-mileage columns -- 20 of 20 and 18 of 18 -- and scores 0 of 34 and 0 of 16 on the two cause columns, which is the whole of its 34.81 pct. Adding the cause table and the prior-repair record is pure code and takes it to 76.30 pct. What is left after that is 22 lines in one column, the combination-operation half of labour time, where every printed figure is internally consistent and no free floor moves off 0 -- and the paid arm takes 22 of 22. Buy a model for that column or for none of them.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The whole adjudication on each claimed line -- disposition, ground, provision, record whether each of the 270 claimed lines got an adjudication you could actually put in front of the dealer appealing it: the disposition the answer key carries, and on a DENY the right ground of the six, the policy provision the pack itself prints for it, every record identifier the ground needs, and nothing cited that is not a provision or record the pack names
$0.00
no
yes
the fast tier 44.4% claim adjudication accuracy · the fast tier, correspondence-blind 45.9% claim adjudication accuracy · strongest free floor 76.3% claim adjudication accuracy · second free floor 70.4% claim adjudication accuracy · the assigned free floor 34.8% claim adjudication accuracy · up to 14 more measured on a run
Verdict only, three-way the same three-way call WITHOUT the citation and record requirements -- what a rules engine can score. The gap between this and the row above is the part of the work a coverage-window check does not do, in numbers: 72.22 against 44.44 on the paid arm, and 76.30 against 76.30 on the free floor, which cites correctly by construction
$0.00
no
yes
the fast tier 44.4% claim adjudication accuracy · the fast tier, correspondence-blind 45.9% claim adjudication accuracy · strongest free floor 76.3% claim adjudication accuracy · second free floor 70.4% claim adjudication accuracy · the assigned free floor 34.8% claim adjudication accuracy · up to 14 more measured on a run
The gate on the answer key itself whether the key everything else is scored against is internally sound: eleven properties, red-proven with four seeded defects that must each be convicted by name and an acquittal that must still pass afterwards
$0.00
no
yes
no headline metric on its single run — it records violations · properties asserted · seeded defects convicted · acquittal after self test
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
⚠︎ THE FREE FLOOR AND THE PAID ARM ARE SEPARABLE ON THIS CORPUS AND THE FLOOR IS AHEAD: 76.30 pct against 44.44 over 270 lines, a 31.86-point gap, and the floor also wins five supporting columns outright. But the two arms are NOT separable in the direction the discriminator suggests, and reading it as 'the model is worse at warranty claims' is the wrong reading. Split the corpus by channel and they are not close on either half. On the 108 lines the printed tables settle, both take 108 of 108 -- indistinguishable, and the floor does it for $0.00. On the 22 lines that only a sentence in the service correspondence settles, the arm takes 22 of 22 and every floor takes 0 of 22 -- and no tolerance, no threshold and no additional pure-code rule moves that zero, because nothing in those notes is decidable by a regular expression. ⛑ THE ABLATION IS WHAT MAKES THAT A MEASUREMENT RATHER THAN AN INTERPRETATION. s001-warranty-claim-notes-blind is the same arm with the correspondence truncated out of the prompt, and five of its columns land on the strongest free floor's own figure to the decimal: correspondence grounds 0 of 22, the prose half of the labour trap 0 of 22, labour time exceeded 47.62 pct, unpayable lines denied 83.08 pct, silent over-denial 42 of 42. An arm that had been pattern-matching on the tables would barely have moved. This one lost exactly the prose cases and kept exactly the table ones. ⚠︎ AND THE ABLATION SEPARATED THE DISCRIMINATOR FROM THE JOB: the blind arm scores HIGHER on it, 45.93 against 44.44, while its verdict accuracy falls 72.22 -> 54.81 and its over-denials rise 62 -> 91 of 124. Measured driver: blinding removes the bulletin, hotline and goodwill numbers it was quoting into a record-identifier field, so lines carrying a non-identifier citation fall 72 -> 23 and WP-3.4 substitutions fall 42 -> 20. A metric with a citation term in it can be improved by giving the arm less to cite. Both figures are published and neither is adjusted.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Every fact that decides a claim on your bench is already a FIELD -- an in-service date, an odometer reading, a cause code in a table with a covered/not-covered column, an operation code in a published time guide, a prior-repair record with an authorisation number on it -- and your dealers claim labour in hundredths of an hour.
the free floor -- claim-gate, $0.00, and set TOL to your own rounding
It scores 76.30 pct on the discriminator against the fast tier's 44.44, cites the provision correctly on 100 pct of its denials against 66.15, names all 16 unsettleable lines against 6, reads the published labour time right on 261 of 261 against 239, and invents nothing. It runs all 45 packs in 0.1 s with no key and no network.
⚠︎ Do not use the free floor where any decisive fact is a sentence. It takes 0 of 22 correspondence grounds and denies all 42 of the lines the pack itself has already answered -- sent as written, that is 42 denials to dealers holding the manufacturer's own bulletin, hotline case or goodwill authorisation. And do not take its 76.30 pct on trust: it reads 57.78 pct at TOL = 0.00, so sweep the constant on your own corpus before believing any of it.
Your bulletins, hotline authorisations and goodwill decisions live in prose -- a technical information system, a CRM case, a region's spreadsheet -- and a claim that looks out of coverage on the printed policy is routinely reopened by one of them.
the paid arm, and budget for the over-denials
This is the only column free code does not reach at any tolerance: 22 of 22 for the fast tier against 0 of 22 for all three floors, and the ablation proves it is real reading -- blind the same arm and it drops to 0 of 22 and lands on the floor's figure in five columns. It also cuts the expensive failure by three and a half times: 12 silent over-denials of 42 against the floor's 42 of 42.
⚠︎ Do not buy the paid arm expecting a defensible adjudication out of the box. It denies 62 of 124 payable lines, every one of them on a rounding of 0.03 hours or less, cites a provision the key does not accept on all 42 labour-time cells, and puts a non-identifier in the record field on 72 of 269 lines. Budget for a tolerance you apply yourself and for a citation post-check; without both, half your denials come back.
You want one number to decide with.
neither -- read the column split
⚠︎ THE SINGLE NUMBER IS THE MOST MISLEADING THING ON THIS PAGE. The assigned floor ties every arm on the date column (20 of 20) and the mileage column (18 of 18) and scores 0 of 34 and 0 of 16 on the two cause columns. The strongest floor closes both cause columns for nothing and still takes 0 of 22 on the combination-operation half of labour time. The paid arm takes that 22 and pays for it with 62 over-denials. No single accuracy figure carries any of that.
⚠︎ Do not read the column split as permission to run both and take the better answer per line. Nothing here measures an ensemble, no arbitration rule was tested, and the two arms disagree on the lines that matter most -- the free floor is confidently wrong on all 42 already-answered lines and the paid arm is confidently wrong on 62 roundings.
You are pointing this at real transmitted claim records out of a dealer management system.
rewrite src/claim.py first, and expect the correspondence column to disappear
src/claim.py parses this corpus's fixed-column layout and returns EMPTY tables against anything else rather than guessing, so every line would come back INSUFFICIENT_EVIDENCE -- an honest failure, not a hidden one. And the half this kit measures as the paid arm's premium is exactly the half that is not in the claim feed at all in a real estate. data/SOURCES.md says what breaks, field by field.
⚠︎ Do not assume the correspondence column survives contact with your estate. On this corpus the sentence that reverses a denial is inside the pack; in a real one it is in a technical information system, a CRM case or a region's spreadsheet, and if you do not put it in front of the model the paid arm's entire 22-of-22 premium disappears -- which is exactly what s001-warranty-claim-notes-blind measured when the section was removed.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
NO_TOLERANCE_TO_TYPE
the dealer's own hundredth-of-an-hour rounding denied as a labour over-claim
62
WC-0001 L1, WC-0002 L1, WC-0002 L6, WC-0004 L3, WC-0004 L5 and 57 more. Every one is 0.01 h (23 lines), 0.02 h (13) or 0.03 h (26) above the published operation. 50 fall on CLEAN_LINE, 7 on the coverage-extension trap and 5 on the goodwill trap.
BETTER_PROVISION_THAN_THE_KEY
cited WP-3.4 where the answer key carries WP-3.1, on every labour-time cell
42
WC-0001 L5, WC-0002 L2, WC-0002 L5, WC-0004 L1, WC-0010 L1 and 37 more. WP-3.1 is "Labour is payable at the time the labour operation guide publishes"; WP-3.4 is "Time above the published operation is payable only with a hotline authorisation". Both are…
REAL_STRING_WRONG_FIELD
cited something printed in the pack that is not an identifier the pack defines
72
WC-0002 L1, WC-0002 L2, WC-0002 L3, WC-0002 L4, WC-0002 L5 and 67 more. The tokens are operation codes (40), special coverage adjustment numbers (14), technical hotline cases (13), repair order numbers (10), goodwill authorisations (5) and the literal phrase…
UNSETTLEABLE_TURNED_INTO_A_CALL
answered a line the pack does not settle instead of saying what is missing
10
WC-0003 L3, WC-0003 L6, WC-0009 L1, WC-0009 L5, WC-0015 L4 and 5 more. Nine are an operation code with no row in the labour time guide and seven are a line with no technician story on file; the arm answered 10 of the 16 anyway. The strongest free floor names…
PAYABLE_CALLED_UNSETTLEABLE
a payable line answered INSUFFICIENT_EVIDENCE
3
WC-0033 L6, WC-0037 L5, WC-0038 L4. The mildest failure on the page and the only one that costs nobody money -- it asks for a document instead of denying.
What we could NOT verify
⚠︎ THE ANSWER KEY IS WRONG ON 42 CELLS BY THE MODEL'S READING AND WAS NOT FIXED. Every labour-time denial cites WP-3.4 where the key carries WP-3.1, and WP-3.4 is arguably the better clause. 44.44 pct is therefore a FLOOR on this arm and nothing higher is claimed; the run was not re-fired and the key was not edited.
The 72 lines carrying a non-identifier citation are scored as wrong and there is a real argument they are half-right: 0 of them are absent from the pack, so the arm quoted a genuine bulletin, hotline case or operation code into a field that wanted a record identifier. Both the strict figure (26.77 pct) and the absent-from-pack figure (0.00 pct) are published so a reader can choose which they believe.
Repeat variance, on every arm. r001, s001 and both calibration probes each ran exactly once, and 96.00 pct of r001's output tokens are provider-side reasoning the provider re-rolls per call. Nothing here says the same prompt returns the same adjudication tomorrow.
A second model. Four arms are scored and three of them are pure code; only one tier was bought. The provider seam is one function in src/adapters/__init__.py and a second tier is .env plus one more run -- that is a spend decision, recorded rather than left to look like a result.
⚠︎ INJECTION. evals/injection.py is written, carries both phrasings verbatim and has NEVER been fired against this kit -- there is no results/eval-x001-warranty-claim-injection.json on disk. This page publishes no resistance rate, because a rate nobody measured is worse than an absence that is named. It is a different question from the ablation: blinding removes the sentence, injection leaves it there and adds one arguing the other way.
A keyword floor over the notes. Grepping the correspondence for "special coverage adjustment" or "combination" would close part of the 22-line gap for nothing, and it was deliberately NOT built, because it would be tuned to the sentences this generator happens to write and would measure the generator rather than the method.
⚠︎ A MEASURED CORPUS WART. Of the 16 special coverage adjustment notes, the limb the denial actually turns on reopens the coverage on 16 of 16 -- but both limbs are derived from where the vehicle stands rather than from the coverage term, so 8 of 16 print the OTHER limb at or below the standard term (5 on months, 3 on miles): an extension that shortens the half nobody is arguing about. The trap is intact and the sentence is untidy. Found after r001 was already in flight; regenerating would have discarded 45 paid calls to tidy a sentence, so it is recorded in tools/build_corpus.py with its fix rather than regenerated away.
⚠︎ THE DISCRIMINATOR ITSELF, CONVICTED BY THIS KIT'S OWN ABLATION. The correspondence-blind arm scores HIGHER on it (45.93 against 44.44) while being strictly worse at the job -- verdict accuracy 72.22 -> 54.81, over-denials 62 -> 91. Blinding removes what it was mis-citing, so the citation term rewards an arm for having less to say. Do not read a 1.49-point difference on this metric as a difference in quality.
Whether the corpus's difficulty resembles a real warranty bench. It is generated, its case mix was chosen by the author, and 130 of 270 lines are unpayable -- a rate no real dealer network would survive. Every rate on this page is against that mix and against nothing else.
Concurrency. r001 ran at 5 workers and s001 at 6, against a key sibling kits were using at the same time. The latency figures are bounds, not clean single-tenant measurements, and no run was taken at one worker to establish the floor.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
3,316.0
18,959.91
153,969 ms
$0.058538
the fast tier, correspondence-blind
3,200.6
16,445.53
142,311 ms
$0.050937
strongest free floor, no model
0
0
0 ms
$0.000000
second free floor, no model
0
0
0 ms
$0.000000
the assigned free floor, no model
0
0
0 ms
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-26. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Answering vs gradingAnswering and grading are two separate bills
100 live calls in total: 45 scored, 45 correspondence-blind, 4 on the discarded 32,000-token probe, 4 on the re-taken 64,000 probe, and 2 at max_tokens=1 for the prompt-token split. The three free floors, the tolerance sweep, the stub and the answer-key gate cost $0.00 and can be re-run by anyone with no key. ⚠︎ THE DISCARDED PROBE'S COST IS LISTED RATHER THAN NETTED OUT: it was paid for, its scores are not published, and hiding it would understate what this kit cost to build.
Cost driversWhat actually moves the bill
THE OUTPUT SIDE, AND IT IS NOT CLOSE. 18,959.91 output tokens per claim against 3,316.00 input -- at this card the output is 96.9 pct of the bill.
REASONING THE PROVIDER RE-ROLLS PER CALL. 818,949 of 853,196 output tokens across the scored run -- 96.00 pct -- are provider-side reasoning rather than the adjudication that gets published. The answer itself is about 760 tokens.
The claim pack, which is two thirds of the input bill: 1,846 of WC-0042's 2,970 input tokens. The instruction is a fixed 1,124 on every call, measured as an intercept rather than counted.
The number of claimed lines on a repair order, indirectly. Every pack here carries exactly six, so this kit CANNOT separate cost-per-claim from cost-per-line and does not pretend to.
Nothing else. There is no index, no retrieval, no chunking, no embedding and no second pass -- one call per claim, and the free floors add $0.00.
Your volumeWhat it costs at your volume
Linear, and this kit says so rather than implying a discount it did not measure. Every claim is independent -- no carried state, no shared index, no cache across calls -- so 450 claims is $26.342 and 4,500 is $263.42 at this card. What does NOT scale linearly is wall clock: 45 claims took 1,503.4 s at 5 concurrent workers against a key sibling kits were using at the same time, and no run was taken at another concurrency to establish what a second worker buys. The free floors scale to nothing: 0.1 s for 45 packs, and they are pure Python over a text file.
Where pricing changes shape
⚠︎ THE 64,000-TOKEN CEILING IS A CLIFF THIS KIT ALREADY WALKED UP TO ONCE. c000-warranty-claim-calibration drew 30,310 output tokens against a 32,000 cap -- 94.7 pct -- so 32,000 was abandoned before any scored arm fired. The scored run's heaviest reply was 30,562, which would have been marginal at the abandoned cap and is 47.8 pct of the published one. You are billed for tokens DRAWN and not for the cap, so raising it costs nothing until a reply actually uses it; a truncated reply is recorded as a failure and stays in the denominator.
Provider-side reasoning is 96.00 pct of the output bill and is not something this kit controls. A provider that reprices reasoning tokens separately, or a tier that reasons less, moves the cost per claim by more than any change to the prompt would.
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
One tier was bought, and it is the cheap one. The question this kit asks is whether a model buys anything at all over pure code on a task where most facts are fields -- and the answer on the discriminator is no, at 44.44 pct against 76.30. Spending more per call to re-ask that question would have been buying a better answer to the wrong half of it; the column that separates is the 22 correspondence lines, and the fast tier already takes 22 of 22 of them. A deliberating tier is .env plus one more run and is recorded as not bought.
Other modelsCost projection
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
149,220input tokens · this run
853,196output tokens
$0.000what it actually cost
45 claims, 270 claimed lines
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$1.054
$1.054
$23.42
2026-09-12
gemini-3-flash
Google
$2.634
$2.634
$58.54
2026-09-18
gemini-3-8-flash
Google
$3.311
$3.311
$73.59
2026-09-18
llama-5
Meta
$3.813
$3.813
$84.72
2026-09-18
claude-haiku-4-5
Anthropic
$4.415
$4.415
$98.12
2026-09-12
grok-4-5
xAI
$5.418
$5.418
$120.39
2026-09-18
grok-4-6
xAI
$5.418
$5.418
$120.39
2026-09-18
claude-sonnet-5
Anthropic
$8.830
$8.830
$196.23
2026-09-12
gemini-3-1-pro
Google
$10.537
$10.537
$234.15
2026-09-18
gpt-5-6-terra
OpenAI
$10.537
$10.537
$234.15
2026-09-12
gpt-5-6-sol
OpenAI
$17.661
$17.661
$392.46
2026-09-12
claude-opus-4-8
Anthropic
$22.076
$22.076
$490.58
2026-09-12
claude-opus-5
Anthropic
$22.076
$22.076
$490.58
2026-09-12
claude-fable-5
Anthropic
$44.152
$44.152
$981.16
2026-09-18
claude-fable-5-1
Anthropic
$44.152
$44.152
$981.16
2026-09-18
gpt-6-astra
OpenAI
$44.152
$44.152
$981.16
2026-09-17
Read this against the numbers above
NOTHING HERE IS A BILL ANYONE PAID. usd_actually_paid is 0.0 because the provider that ran this publishes no rate card this repo commits.
Every row prices THIS kit's measured token counts against ANOTHER model's published rate. A model that reasons more or less than the one measured will not produce these tokens.
96.00 pct of the output tokens being priced are provider-side reasoning. A provider that bills reasoning differently invalidates every row.
The free floors are $0.00 and beat the measured arm on the discriminator, so the cheapest row on this page is also the most accurate one.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
16 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Nine of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/segment.pysection split
Cut the claim pack on its underlined headings -- Warranty Claim Review, Warranty Policy And Guides, Claimed Lines, Records On File, Service Correspondence, Dealer Position. One regular expression, deterministic, so "what did we send" is a list of six section names and not an argument about a token window. A pack with no headings comes back as one unnamed section rather than raising, so a forker pointing this at their own export gets a usable failure instead of a stack trace.
src/segment.py
# Cut a warranty claim review pack into addressable sections. Pure code, no model.
HEAD = re.compile(r"^([A-Z][A-Za-z ,]+)\n-{10,}\n", re.M)
def sections(text):
def named(text, name):
src/select.pythe seam that withholds — a swap seam
The Dealer Position block -- the store's open chargeback exposure, the goodwill amount a district manager may pay without a region review, and the manufacturer's own relationship posture toward this dealer -- never leaves the machine. None of it bears on whether a cause code is a covered defect, whether the vehicle was inside its coverage on the repair date, or whether the hours claimed match the published operation. All of it is a reason to pay. Withheld at the seam rather than trusted to a sentence in the instruction, and the UI prints, per claim, which sections went and which stayed.
You change it to: NEVER_SENT. One tuple, holding one section name. It is a named-section denylist and not a redaction system: the same goodwill authority that is withheld in the Dealer Position block will be sent when a sentence in the Service Correspondence names it, because the correspondence is where the answers live.
src/select.py
# SEAM 2 -- what leaves this machine. Pure code, and deliberately a denylist of one.
NEVER_SENT = ("Dealer Position",)
def sent(sec_names):
def body(text, sections_fn):
src/claim.pythe pack parser
The coverage terms table with its months and miles columns, the labour operation guide with its coverage code and published time, the cause classification table with its covered / wear / maintenance / customer-damage column, the provision index, every claimed line with its operation code, cause code, hours, rate and amounts, and every record on file with its type, reference and authorisation -- all read with regular expressions. The model is never asked for a figure that sits at a fixed offset. months_between() counts COMPLETED months rather than dividing days by 30.44, because a coverage term expires on the same day of the month it started and this corpus deliberately puts repairs within a week of that boundary.
src/checks.pythe structured warranty checks -- and the labour-hour tolerance — a swap seam
The six grounds the pack's OWN TABLES can decide -- NOT_A_COVERED_DEFECT, CUSTOMER_DAMAGE, PRIOR_UNAUTHORISED_REPAIR, OUT_OF_COVERAGE_TIME, OUT_OF_COVERAGE_MILEAGE, LABOUR_TIME_EXCEEDED -- in a fixed PRIORITY order, plus the two shapes that settle nothing (an operation code with no guide row, a line with no technician story on file) which return INSUFFICIENT_EVIDENCE with the missing item named. recompute() answers the use case's second half on EVERY line whether or not anything is wrong with it: the hours the guide publishes, the labour amount those hours and the claim's rate give, and the variance against what was claimed. TOL = 0.05 HOURS is the single published constant that decides how much over is over, and MILE_TOL is fixed at 0 miles and named separately because a coverage limit is a number in a policy and not a measurement with a rounding error.
You change it to: SUPPORT_TYPE, one dict, six entries. This corpus files exactly one record type per ground; a real adjudication attaches several documents to one ground and some of them are images, so this is the mapping to widen first.
src/checks.py
# The structured warranty checks, in pure code. No model, no service correspondence.
TOL = 0.05
MILE_TOL = 0
PRIORITY = ("NOT_A_COVERED_DEFECT", "CUSTOMER_DAMAGE", "PRIOR_UNAUTHORISED_REPAIR",
SUPPORT_TYPE = {
NO_AUTHORISATION = "none"
def recompute(line, parsed):
def unauthorised_prior(line, parsed):
def over_time(line, parsed, tol=TOL):
def _rid(parsed, line, rtype):
src/prompt.pythe prompt, and the correspondence-blind variant — a swap seam
The whole instruction in one place, in send order. It names the three dispositions and the six grounds, because scoring an arm on a vocabulary it was never given measures the wording rather than the arm. What it deliberately does NOT say is where to look -- nothing in it mentions reading the correspondence against the tables, and nothing mentions that a published operation time can be superseded by a combination time. _strip_notes() is the ablation: it TRUNCATES at the Service Correspondence heading and RAISES if that heading is absent, because a silent no-op there would publish an ablation arm that was never ablated.
You change it to:build(text, blind=True), which truncates the Service Correspondence and says so in the instruction. It is the only evidence this kit can offer that an arm READ the prose rather than learning that transmitted warranty claims are over-claimed: take the notes away and an arm that was reading them loses exactly the correspondence cases.
src/prompt.py
# The assembled prompt, in one place, in the order it is sent.
NOTES_HEADING = "Service Correspondence"
SYSTEM = """You are adjudicating a dealer warranty claim for a vehicle manufacturer.
BLIND_SYSTEM_SUFFIX = """
def build(text, blind=False):
def _strip_notes(body):
def render(parts):
src/draft.pythe model call and the reply parse
One claim pack in, one drafted adjudication out, and the only place a model is called for a claim. Holds the 64,000-token output ceiling that this kit's own probe forced: c000-warranty-claim-calibration fired the four heaviest packs at 32,000 and drew 30,310 output tokens on the largest reply -- 94.7 per cent of that cap -- so 32,000 was abandoned BEFORE any scored arm fired rather than after one came back truncated. c001-warranty-claim-calibration re-took the probe at 64,000 on the corrected corpus and the same four packs drew a largest reply of 27,045. The scored run then bore the decision out and bounded it: across all 45 claims r001-warranty-claim came back 45 of 45 answered, 0 failures, 0 truncated and 0 lines omitted, with its heaviest reply at 30,562 output tokens -- 47.8 per cent of the raised ceiling and ABOVE the cap that was abandoned. The parse is tolerant of a fenced block and of leading prose, uppercases only the closed vocabularies, and coerces a single record identifier returned as a bare string into a list.
src/draft.py
# One claim pack in, one drafted adjudication out. The only place a model is called for a claim.
MAX_TOKENS = 64000
THINKING = None
def parse_reply(text):
def normalise(obj):
def draft(cfg, text, blind=False, complete_fn=None, max_tokens=None):
def cited_ids(answer):
def action_from_lines(answer):
src/adapters/__init__.pythe adapter — a swap seam
One interface, several providers, raw HTTP and no vendor SDK, so the kit installs nothing to call whichever key a forker already holds. ⚑ ITS SOCKET TIMEOUT AND THE OUTPUT CEILING ARE ONE SETTING WEARING TWO NAMES: TIMEOUT_S is 1,800 seconds beside a 64,000-token ceiling, because completions here are not streamed and a reply that actually fills that ceiling holds a silent socket for the whole generation. Transport retries are cut to ONE against four for a transient status, on the ground that a blown socket on a non-streamed generation is not a rate limit and asking four more times turns one defect into five bills.
You change it to: PROVIDER, BASE_URL and MODEL in the repo-root .env, or a kit-local .env carrying MODEL alone to point this one kit at a different tier while keeping the shared account. Adding a provider is one function and one entry in PROVIDERS, and it must return token counts as well as text -- an adapter that returns no usage cannot be published here, because the cost lens prices what the eval lens counted.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
TIMEOUT_S = 1800
TRANSPORT_RETRIES = 1
def _post(url, headers, payload, timeout=TIMEOUT_S):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
src/config.pythe shared connection, and the kit-local override
Provider, base URL, key and model are read from the repo-root .env first, then this kit's own .env, then the real environment, MERGED KEY BY KEY rather than file by file -- so a kit-local file holding MODEL alone points one kit at a different tier and keeps the shared account. Neither file is required and a missing one is not an error, which is what keeps a single cloned kit runnable with no repo above it. sources() reports WHICH file contributed and never a value, so "where is this key coming from" is answerable without anyone reading a key aloud.
src/config.py
# Read .env. No dependency, and no key ever leaves this machine.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
ENV = os.path.join(HERE, ".env")
SHARED_ENV = os.path.join(os.path.dirname(os.path.dirname(HERE)), ".env")
VARS = ("PROVIDER", "BASE_URL", "API_KEY", "MODEL", "EMBED_MODEL")
def _read(path):
def load():
def sources():
MODEL_DISPLAY = "the fast tier"
def has_key(cfg=None):
src/budget.pythe call cap
One key funds every kit under this root, which is the point of the shared .env and also the hole in it. This counts CALLS, not dollars, because a dollar cap needs a rate card the kit does not have and a cap that silently reads the wrong card is worse than none. The ledger line is written BEFORE the call, so a crash mid-call over-counts by one rather than under-counting. No cap configured means no cap and it says so out loud in the line printed before a run; there is no hard-coded ceiling in the file.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
evals/baseline.pythe three free floors — a swap seam
coverage-window is the strawman that is not a strawman -- the months-and-miles arithmetic plus a guide lookup, which is what a warranty system already does the moment a claim is transmitted. line-sweep runs the six structured checks and cites the provision the policy table prints beside the row it used, but never stops: anything it cannot settle it denies anyway. claim-gate is line-sweep plus the one thing that separates a defensible adjudication from a list of exceptions -- it REFUSES to state a ground it cannot support, and every provision and record identifier it cites is read out of the pack rather than assembled from a constant. ⚑ AND THE THIRD FLOOR WINS THE DISCRIMINATOR OUTRIGHT, MEASURED: b002-warranty-claim-claimgate scores 76.30 per cent against r001-warranty-claim's 44.44 per cent over the same 270 lines, free, offline and in 0.1 seconds. The three floors read 34.81, 70.37 and 76.30 per cent respectively. All three are correspondence-blind BY CONSTRUCTION -- 0 of 22 correspondence grounds against the paid arm's 22 of 22 -- and a keyword floor over the notes would close part of that gap and was deliberately not built, because it would be tuned to the sentences this generator happens to write and would measure the generator.
You change it to: --floor coverage-window | line-sweep | claim-gate. Three pure-code adjudicators behind the same answer shape, so one scorer grades all three and the model without a second code path.
evals/baseline.py
# THREE FREE FLOORS. No key, no model, no network. Each is a genuine attempt at the job.
MODES = ("coverage-window", "line-sweep", "claim-gate")
def _action(rows):
def _row(line, disposition, ground, provision, support, finding, allowed=None):
def review(text, mode="claim-gate", tol=checks.TOL):
def _window(line, p, rc, tol):
evals/tolerance.pythe tolerance sweep
Three floors by eight labour-hour tolerances, free and offline, and it says the headline is a measurement of a parameter as much as of a method. claim-gate scores 57.78 per cent at 0.00 hours, 76.30 at 0.05, holds 76.30 at 0.10 and falls to 74.81 at 1.00. The weakest floor does the opposite and climbs to its sweep best of 39.26 per cent at a one-hour tolerance while its labour catch collapses from 47.62 to 9.52 -- it is not getting better, it is buying accuracy by giving up the only thing it was looking for. No floor is published outside the band over which the answer key does not move, because a floor allowed to relabel the key is not a floor.
evals/tolerance.py
# THE TOLERANCE SWEEP. Free, offline, and it must run before anyone believes a floor's score.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
RESULTS = os.path.join(HERE, "results")
TOLERANCES = (0.00, 0.005, 0.02, 0.05, 0.10, 0.25, 0.50, 1.00)
def _docs():
def _text(d):
def sweep():
def main():
evals/scoring.pythe scorer
Exact match per line, no model grades anything, and every rate carries its own denominator. A line counts toward claim_adjudication_accuracy_pct only when the disposition is right AND, on a denial, the ground is the right one of six, the provision cited is the one the pack carries for it, every required record identifier is attached and nothing on the line was invented. The catch rate is split by CHANNEL -- table grounds against correspondence grounds -- and labour time is split again on top of that, because the named trap is planted on both channels and a blended number would let one half carry the other. silent_overdeny_pct and invented_citation_pct are the two named dangerous failures, and the majority-class score is printed beside every arm.
evals/scoring.py
# Score an arm against the answer key. Pure code, exact match per line. No model grades anything.
DENY, PAYABLE, INSUFFICIENT = "DENY", "PAYABLE", "INSUFFICIENT_EVIDENCE"
OMITTED = "OMITTED"
HOUR = 0.005
def _pct(n, d):
def _key(s):
def rows_by_line(answer):
def _cited(r):
def _numeric(v):
def score(records, golds, valid_ids=None, allowed_hours=None, texts=None):
evals/check_labels.pythe answer-key gate
Eleven properties, re-derived from the shipped corpus, and the three that matter are the corpus's claims about itself rather than its schema: the table half must re-derive EXACTLY from src/checks.structured_finding(); a correspondence denial must look PAYABLE to the tables; a correspondence trap must look DENIED to them. A prose denial a table can already see is not a prose case and labelling it one would inflate the very column this kit exists to measure. --self-test red-proves the gate against seeded defects rather than asserting that it works.
evals/check_labels.py
# A GATE ON THE ANSWER KEY ITSELF. Free, offline, no model. Run it before believing any arm.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
TABLE_CASES = {"S_NOT_A_COVERED_DEFECT", "S_CUSTOMER_DAMAGE", "S_PRIOR_UNAUTHORISED_REPAIR",
PROSE_DENY_CASES = {"C_LABOUR_TIME_EXCEEDED"}
PROSE_TRAP_CASES = {"P_COVERAGE_EXTENDED", "P_TIME_AUTHORISED", "P_REPAIR_AUTHORISED"}
DERIVABLE = TABLE_CASES | {"CLEAN_LINE", "M_OP_NOT_IN_GUIDE", "M_NO_TECHNICIAN_STORY"}
def load_gold():
evals/injection.pythe injection probe — a swap seam
One instruction-shaped note forced into the Service Correspondence of every claim in a scope FIXED BEFORE THE RUN -- the claims where the paired run correctly denied at least one line, in claim-id order, capped by an argument that is recorded in the result file. Every injected cell is paired against its own un-injected answer rather than against gold, because a cell the arm was already getting wrong cannot be suppressed. It records whether a cell that HELD had its ground, its provision or its support changed underneath it, and both phrasings are quoted in the file in full, including any that was not fired.
You change it to: --phrasing, over the PHRASINGS dict, with --scope and --paired-with fixing the denominator before anything is sent. Sibling kits have measured one phrasing at over eighty per cent suppression and another at near zero, so the phrasings are reported one by one and never meaned together.
evals/injection.py
# THE INJECTION PROBE. Does an instruction-shaped note in the service correspondence talk the
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
KIT = "warranty-claim"
HEADING = "Service Correspondence"
PHRASINGS = {
def documents():
def load_doc(doc_id):
src/app.pythe local UI
http.server, standard library only, five endpoints and no write path -- nothing here pays, denies, submits or charges back a claim, and there is no endpoint, button or flag that adds one. It renders with no key: the pack, everything read off it in code, the published labour time on every line and the strongest free floor all cost nothing, so the page always computes them. The floor's adjudication sits beside the model's with an Agree? column, so a line where the two agree is a line where the model bought nothing and the page says so rather than taking the credit.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
CORPUS = os.path.join(HERE, "data", "corpus")
PORT = int(os.environ.get("PORT", "9074"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-warranty-claim")
FLOOR_MODE = "claim-gate"
def documents():
def load_doc(doc_id):
def read_in_code(text):
tools/build_corpus.pythe corpus generator — a swap seam
Writes data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json from seed 20260826. The answer key is DERIVED for everything pure code can settle -- it is structured_finding()'s own output -- so the key and the free floor cannot silently drift apart. What is TYPED is only the four cases code cannot reach, and each of those is then asserted to look the opposite way to the tables. The dealer's own rounding is planted at no more than 0.03 hours and every planted over-claim at 0.08 hours or more, which is the gap that makes the tolerance sweep readable. ⚠︎ IT ALSO CARRIES A RECORDED DEFECT AND ITS FIX RATHER THAN A REGENERATED CORPUS: both limbs of a special coverage adjustment note are derived from where the vehicle stands rather than from the coverage term, so 8 of the 16 notes print the limb nobody is arguing about at or below the standard term. It was found after r001-warranty-claim was already in flight, and regenerating would have discarded 45 paid calls to tidy a sentence.
You change it to: One seed (20260826) and one case plan. Replace it with your own claim packs and your own data/gold.jsonl; data/SOURCES.md lists what to change and in what order, and python3 -m evals.check_labels is what tells you the replacement key is self-consistent.
tools/build_corpus.py
# GENERATE the shippable corpus: 45 warranty claim review packs, an answer key, and the stats.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260826
PACKS = 45
LINES = 6
DATASET = "warranty-claim-2026-08-26-45packs-v2"
RULE = "-" * 64
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
src/segment.pyCut the claim pack on its underlined headings -- Warranty Claim Review, Warranty Policy And Guides, Claimed Lines, Records On File, Service Correspondence, Dealer Position. One regular expression, deterministic, so "what did we send" is a list of six section names and not an argument about a token window. A pack with no headings comes back as one unnamed section rather than raising, so a forker pointing this at their own export gets a usable failure instead of a stack trace.
src/select.pyThe Dealer Position block -- the store's open chargeback exposure, the goodwill amount a district manager may pay without a region review, and the manufacturer's own relationship posture toward this dealer -- never leaves the machine. None of it bears on whether a cause code is a covered defect, whether the vehicle was inside its coverage on the repair date, or whether the hours claimed match the published operation. All of it is a reason to pay. Withheld at the seam rather than trusted to a sentence in the instruction, and the UI prints, per claim, which sections went and which stayed. A swap seam.
src/claim.pyThe coverage terms table with its months and miles columns, the labour operation guide with its coverage code and published time, the cause classification table with its covered / wear / maintenance / customer-damage column, the provision index, every claimed line with its operation code, cause code, hours, rate and amounts, and every record on file with its type, reference and authorisation -- all read with regular expressions. The model is never asked for a figure that sits at a fixed offset. months_between() counts COMPLETED months rather than dividing days by 30.44, because a coverage term expires on the same day of the month it started and this corpus deliberately puts repairs within a week of that boundary.
src/checks.pyThe six grounds the pack's OWN TABLES can decide -- NOT_A_COVERED_DEFECT, CUSTOMER_DAMAGE, PRIOR_UNAUTHORISED_REPAIR, OUT_OF_COVERAGE_TIME, OUT_OF_COVERAGE_MILEAGE, LABOUR_TIME_EXCEEDED -- in a fixed PRIORITY order, plus the two shapes that settle nothing (an operation code with no guide row, a line with no technician story on file) which return INSUFFICIENT_EVIDENCE with the missing item named. recompute() answers the use case's second half on EVERY line whether or not anything is wrong with it: the hours the guide publishes, the labour amount those hours and the claim's rate give, and the variance against what was claimed. TOL = 0.05 HOURS is the single published constant that decides how much over is over, and MILE_TOL is fixed at 0 miles and named separately because a coverage limit is a number in a policy and not a measurement with a rounding error. A swap seam.
src/prompt.pyThe whole instruction in one place, in send order. It names the three dispositions and the six grounds, because scoring an arm on a vocabulary it was never given measures the wording rather than the arm. What it deliberately does NOT say is where to look -- nothing in it mentions reading the correspondence against the tables, and nothing mentions that a published operation time can be superseded by a combination time. _strip_notes() is the ablation: it TRUNCATES at the Service Correspondence heading and RAISES if that heading is absent, because a silent no-op there would publish an ablation arm that was never ablated. A swap seam.
src/draft.pyOne claim pack in, one drafted adjudication out, and the only place a model is called for a claim. Holds the 64,000-token output ceiling that this kit's own probe forced: c000-warranty-claim-calibration fired the four heaviest packs at 32,000 and drew 30,310 output tokens on the largest reply -- 94.7 per cent of that cap -- so 32,000 was abandoned BEFORE any scored arm fired rather than after one came back truncated. c001-warranty-claim-calibration re-took the probe at 64,000 on the corrected corpus and the same four packs drew a largest reply of 27,045. The scored run then bore the decision out and bounded it: across all 45 claims r001-warranty-claim came back 45 of 45 answered, 0 failures, 0 truncated and 0 lines omitted, with its heaviest reply at 30,562 output tokens -- 47.8 per cent of the raised ceiling and ABOVE the cap that was abandoned. The parse is tolerant of a fenced block and of leading prose, uppercases only the closed vocabularies, and coerces a single record identifier returned as a bare string into a list.
src/adapters/__init__.pyOne interface, several providers, raw HTTP and no vendor SDK, so the kit installs nothing to call whichever key a forker already holds. ⚑ ITS SOCKET TIMEOUT AND THE OUTPUT CEILING ARE ONE SETTING WEARING TWO NAMES: TIMEOUT_S is 1,800 seconds beside a 64,000-token ceiling, because completions here are not streamed and a reply that actually fills that ceiling holds a silent socket for the whole generation. Transport retries are cut to ONE against four for a transient status, on the ground that a blown socket on a non-streamed generation is not a rate limit and asking four more times turns one defect into five bills. A swap seam.
src/config.pyProvider, base URL, key and model are read from the repo-root .env first, then this kit's own .env, then the real environment, MERGED KEY BY KEY rather than file by file -- so a kit-local file holding MODEL alone points one kit at a different tier and keeps the shared account. Neither file is required and a missing one is not an error, which is what keeps a single cloned kit runnable with no repo above it. sources() reports WHICH file contributed and never a value, so "where is this key coming from" is answerable without anyone reading a key aloud.
src/budget.pyOne key funds every kit under this root, which is the point of the shared .env and also the hole in it. This counts CALLS, not dollars, because a dollar cap needs a rate card the kit does not have and a cap that silently reads the wrong card is worse than none. The ledger line is written BEFORE the call, so a crash mid-call over-counts by one rather than under-counting. No cap configured means no cap and it says so out loud in the line printed before a run; there is no hard-coded ceiling in the file.
evals/baseline.pycoverage-window is the strawman that is not a strawman -- the months-and-miles arithmetic plus a guide lookup, which is what a warranty system already does the moment a claim is transmitted. line-sweep runs the six structured checks and cites the provision the policy table prints beside the row it used, but never stops: anything it cannot settle it denies anyway. claim-gate is line-sweep plus the one thing that separates a defensible adjudication from a list of exceptions -- it REFUSES to state a ground it cannot support, and every provision and record identifier it cites is read out of the pack rather than assembled from a constant. ⚑ AND THE THIRD FLOOR WINS THE DISCRIMINATOR OUTRIGHT, MEASURED: b002-warranty-claim-claimgate scores 76.30 per cent against r001-warranty-claim's 44.44 per cent over the same 270 lines, free, offline and in 0.1 seconds. The three floors read 34.81, 70.37 and 76.30 per cent respectively. All three are correspondence-blind BY CONSTRUCTION -- 0 of 22 correspondence grounds against the paid arm's 22 of 22 -- and a keyword floor over the notes would close part of that gap and was deliberately not built, because it would be tuned to the sentences this generator happens to write and would measure the generator. A swap seam.
evals/tolerance.pyThree floors by eight labour-hour tolerances, free and offline, and it says the headline is a measurement of a parameter as much as of a method. claim-gate scores 57.78 per cent at 0.00 hours, 76.30 at 0.05, holds 76.30 at 0.10 and falls to 74.81 at 1.00. The weakest floor does the opposite and climbs to its sweep best of 39.26 per cent at a one-hour tolerance while its labour catch collapses from 47.62 to 9.52 -- it is not getting better, it is buying accuracy by giving up the only thing it was looking for. No floor is published outside the band over which the answer key does not move, because a floor allowed to relabel the key is not a floor.
evals/scoring.pyExact match per line, no model grades anything, and every rate carries its own denominator. A line counts toward claim_adjudication_accuracy_pct only when the disposition is right AND, on a denial, the ground is the right one of six, the provision cited is the one the pack carries for it, every required record identifier is attached and nothing on the line was invented. The catch rate is split by CHANNEL -- table grounds against correspondence grounds -- and labour time is split again on top of that, because the named trap is planted on both channels and a blended number would let one half carry the other. silent_overdeny_pct and invented_citation_pct are the two named dangerous failures, and the majority-class score is printed beside every arm.
evals/check_labels.pyEleven properties, re-derived from the shipped corpus, and the three that matter are the corpus's claims about itself rather than its schema: the table half must re-derive EXACTLY from src/checks.structured_finding(); a correspondence denial must look PAYABLE to the tables; a correspondence trap must look DENIED to them. A prose denial a table can already see is not a prose case and labelling it one would inflate the very column this kit exists to measure. --self-test red-proves the gate against seeded defects rather than asserting that it works.
evals/injection.pyOne instruction-shaped note forced into the Service Correspondence of every claim in a scope FIXED BEFORE THE RUN -- the claims where the paired run correctly denied at least one line, in claim-id order, capped by an argument that is recorded in the result file. Every injected cell is paired against its own un-injected answer rather than against gold, because a cell the arm was already getting wrong cannot be suppressed. It records whether a cell that HELD had its ground, its provision or its support changed underneath it, and both phrasings are quoted in the file in full, including any that was not fired. A swap seam.
tools/build_corpus.pyWrites data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json from seed 20260826. The answer key is DERIVED for everything pure code can settle -- it is structured_finding()'s own output -- so the key and the free floor cannot silently drift apart. What is TYPED is only the four cases code cannot reach, and each of those is then asserted to look the opposite way to the tables. The dealer's own rounding is planted at no more than 0.03 hours and every planted over-claim at 0.08 hours or more, which is the gap that makes the tolerance sweep readable. ⚠︎ IT ALSO CARRIES A RECORDED DEFECT AND ITS FIX RATHER THAN A REGENERATED CORPUS: both limbs of a special coverage adjustment note are derived from where the vehicle stands rather than from the coverage term, so 8 of the 16 notes print the limb nobody is arguing about at or below the standard term. It was found after r001-warranty-claim was already in flight, and regenerating would have discarded 45 paid calls to tidy a sentence. A swap seam.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 3316 input and 18959 output tokens per claim (one transmitted warranty claim, covering every line claimed on it), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Claim (one transmitted warranty claim, covering every line claimed on it)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per claim (one transmitted warranty claim, covering every line claimed on it) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
NOT ATTACKED YET, AND THE PAGE SAYS SO RATHER THAN LOOKING FULL. One boundary on this kit IS measured, and it is the one the use case turns on: the Dealer Position section -- the dealer's open chargeback exposure, the goodwill amount the district manager may authorise without a region review, and the manufacturer's relationship posture toward the store -- never leaves this machine. It is withheld at the seam in pure code (src/select.NEVER_SENT), on 45 of 45 packs, with sections_used recorded as ['system', 'pack'] on every call of the only live run. None of it bears on whether a cause is a covered defect, whether the vehicle was inside coverage on the repair date, or whether the hours claimed match the published operation time; all of it is exactly what you would least like a generated adjudication to have been reasoning about. THE MODEL IS NEVER HANDED A REASON TO PAY. The boundary that is NOT measured is the one that carries the answers: the Service Correspondence reaches the model verbatim, it is the only evidence for 64 of the 270 lines, and no injection probe has been fired against this kit.
The credential is looked for in three places and the most specific one wins: the repository-wide <code>.env</code>, then an optional file belonging to this kit alone, then whatever the real environment holds. Merging happens variable by variable, so a kit-level file naming only MODEL redirects this one adjudicator while continuing to use the shared account. Neither file has ever been tracked — both were ignored from the first commit and no credential has existed in this repository. <code>src/config.sources()</code> answers where a setting came from by naming files and never contents, so that question can be settled without anybody reading a secret out loud. <code>save()</code> creates the file already at 0600, so there is no window at the default umask in which it sits readable. If the provider returns an error, the local UI strips the key and the base URL out of the message before it reaches the browser. A reader is never asked for any of this: the interface loads, browses, replays and computes floors with nothing configured, which on warranty claims is not a crippled mode at all, since the best-scoring arm the kit publishes calls no one.
The experimentWe attacked it — two gates
An indirect prompt injection has to clear two gates, and they behave differently. Both were run for real on 2026-08-26.
Boundary checked
What could go wrong
What the code guarantees
Whether the dealer's own commercial position can reach the model
A sentence in the prompt telling the model to disregard the dealer's position. That is a request, not a boundary, and it is unfalsifiable from the outside -- you cannot tell a model that ignored a section from a model that read it and did not say so.
MEASURED, AND IT HOLDS BY CONSTRUCTION. src/select.py removes the named section before the pack is assembled, so the text is never in the request body. 45 of 45 packs; sections_used = ['system', 'pack'] on all four calls of c001-warranty-claim-calibration; the UI prints what went and what stayed per claim so a reader can check rather than believe. It is a denylist of ONE named heading and not a redaction system -- a claim mentioning the goodwill authority inside the Service Correspondence sends that mention.
Whether a sentence in the Service Correspondence can withdraw a denial ground the tables already prove
Counting where the generator happened to place such a sentence. That gives a denominator of nothing and would publish a reassuring rate off two or three cells.
NOT RUN. evals/injection.py exists and is complete: it REPLACES the whole Service Correspondence section in every claim in scope, fixes scope BEFORE reading any answer (the claims a named paired run correctly denied at least one line in, in claim-id order, capped as a recorded spend decision), pairs every cell against that same arm's own un-injected answer rather than against the key, scores the whole answer rather than the verdict, and reports the rate, its denominator and its ceiling. Zero calls have been made. There is no results/eval-x001-warranty-claim-injection.json on disk, and the kit publishes no suppression rate -- not a low one.
Whether a citation can be invented
Pointing at the instruction the arm is sent, which tells it to name only provisions and records printed in front of it, and calling that the control.
SCORED BUT NOT PREVENTED, AND ONLY ON FREE CODE. invented_citation_pct and absent_citation_pct are computed against the pack's own identifier set on every arm, and both are 0.0 pct over 150 citing lines on the strongest free floor -- which is structural, because src/checks._prov returns a provision only when the pack prints it. Nothing filters a paid arm's citations, and the valid identifier set is already computed in code and not used as a filter.
Two of these three boundaries are measured and the third is the one that matters most. The withholding seam is verifiable from the request body itself, and citation validity is scored on every run -- but both are properties this kit controls in code. The boundary it does NOT control is the untrusted channel it must read to do the job at all, and that one has never been probed here. A threat model that reported only the two it can prove would be a threat model shaped by what was easy to measure.
The resultNothing has attacked this kit yet -- the boundary that IS measured is the withheld Dealer Position block, gone from 45 of 45 packs, and the boundary that is NOT is the correspondence channel that decides 64 of the 270 lines
0attack calls made against this kit
2phrasings written verbatim, neither fired
No red-team run has been taken on warranty-claim. evals/injection.py is written, tested in shape against a sibling kit's design and never executed here: it would force an instruction-shaped note into every claim in scope, fix scope before reading any answer, pair every cell against the same arm's own un-injected reply so that a cell the arm was already getting wrong cannot be counted as suppressed, and publish the rate with both its denominator and its ceiling. Both phrasings -- one asserting the coverage position was agreed at zone level and closed, one a district-manager directive to raise only certain denials during a retention review -- are printed in full in that file so a reader can re-take the probe rather than take a description on trust. Neither has been sent. The only live calls this kit has ever made are the four of c001-warranty-claim-calibration, fired to size the output-token ceiling.
Read this twice
Every word of the Service Correspondence is put in front of the model exactly as the pack prints it, and that is a decision rather than an oversight. Take that section away and 64 of these 270 lines lose the only thing that decides them — 22 denial grounds nothing else reveals, plus 42 lines where a special coverage adjustment, a technical hotline authorisation or a district manager's goodwill authorisation has already answered the defect the tables think they see. So the section an attacker would write into is the same section a correct adjudication depends on, and no filter separates the two. On this kit nobody has tested what happens when someone does. There is no probe result to quote, which means the risk here is unknown, not small. Pure code shrugs the whole question off, and only because it never opens the section — the identical deafness that leaves it at 0 of 22 correspondence grounds and has it deny all 42 lines the manufacturer's own paperwork already covers.
HonestyWhat this does not prove
WHETHER THIS KIT CAN BE TALKED OUT OF A DENIAL AT ALL. No injection run has been fired. Sibling kits on this estate have measured suppression rates from near zero to over eighty per cent on a single replaced section, and a single number would bound nothing even if one existed here -- but there is not one, and the absence is the finding.
Whether either written phrasing would work. Both are quoted in full in evals/injection.py and neither was sent. Naming them is not measuring them.
Whether a correspondence ground can be SUPPRESSED as opposed to blinded. Replacing the Service Correspondence destroys the evidence those 22 grounds rest on, so even once the probe runs, that channel would have to be reported separately rather than folded into a clean suppression rate.
Whether an instruction hidden in the Dealer Position section would do anything. That text is removed before a request is assembled, so it has no path to a model — and no one has tried to find one.
Whether a paid arm invents citations at any rate worth publishing. The only live run on this kit is four calls over 24 lines, which is a calibration probe and not a measurement of citation discipline.
Anything about the provider side of the boundary. What leaves this machine is enumerated and checked; what the configured provider retains, logs or trains on is the reader's contract with that provider and nothing here tests it.
Whether the withholding seam survives a real claim pack. It matches an underlined heading exactly; a real transmitted claim record has no such heading, src/segment.py returns the whole file as one section, and the seam would then withhold nothing at all. That is the first thing to check before pointing this at anything real.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
It never pays, denies, submits, adjusts or charges back a warranty claim, and it never contacts a dealer. What it produces is an ADJUDICATION a warranty analyst reads before deciding what to do with the claim. INSUFFICIENT_EVIDENCE with the missing item named is part of that adjudication rather than a failure of it -- an operation code with no row in the labour time guide and a line with no technician story are both things the pack does not settle, and saying so is the honest answer. Paying or denying is the manufacturer's act under a dealer agreement the pack references and does not enforce.
Stated in the system prompt on every call, in the README, and on the local UI -- and enforced by there being no such endpoint to remove. src/app.py serves four read paths (/api/claims, /api/prompt, /api/recorded, /api/claim) plus /api/draft, and nothing that mutates anything outside results/. There is no dealer address, no claim submission path and no configuration flag that adds one.
EvidenceDoes it hold?
What
Measured
No write path exists
<code>src/app.py</code> exposes four read paths and one drafting path. Nothing it serves writes anywhere except <code>results/</code>. Asked to draft with no credential present, it answers with a plain sentence explaining that no call was made, instead of an error.
The Dealer Position block never leaves this machine
1 named section held back on 45 of 45 packs (src/select.NEVER_SENT), and sections_used is recorded as ['system', 'pack'] on every one of r001-warranty-claim's 45 calls as well as all four calibration calls. The interface prints what travelled and what stayed for each claim, so this is checkable rather than assertable.
Nothing is published until the key has been rebuilt out of the corpus and checked against itself
11 asserted properties over 270 lines, 0 violations; red-proven by seeding four defects in memory -- a table labour over-claim relabelled as a correspondence one, a payable line relabelled DENY, a provision the pack does not print, and a reverse trap relabelled a real denial -- and convicting all four by name. An acquittal is proved too, because a gate that convicts everything convicts nothing.
Free code cannot cite what the pack does not print -- and the paid arm can
0 of 150 citing lines carry a non-identifier on b002-warranty-claim-claimgate, which is structural: src/checks._prov hands back a provision only where the pack prints one. On r001-warranty-claim it is 72 of 269 (26.77 pct). NOTHING WAS FABRICATED -- 0 of 269 cited a string absent from the pack -- but 89 real strings were quoted into a field that wanted a pullable record.
The spend cap is checked in the adapter, before the request is built
one shared append-only ledger under the repo root, written BEFORE each call so a crash mid-call over-counts by one rather than under-counting. An unparseable cap is announced out loud as NO CAP rather than treated as zero.
A raised ceiling is raised together with the socket timeout, or not at all
MAX_TOKENS 64,000 in src/draft.py against TIMEOUT_S 1,800 in src/adapters. The discarded probe drew 30,310 of 32,000 (94.7 pct), so 32,000 was abandoned before a scored arm ever fired; r001-warranty-claim then ran 45 calls with a largest reply of 30,562 (47.8 pct of the raised cap), 0 failures and 0 truncations.
THE ONE GUARDRAIL THAT MATTERS MOST HAS NOT BEEN TESTED ON THIS KIT
0 injection calls. evals/injection.py is written, carries two phrasings verbatim, forces the condition into every claim in scope and pairs every cell against the same arm's own un-injected answer -- and it has never been fired here. There is no results/eval-x001-warranty-claim-injection.json on disk. The Service Correspondence is the only evidence for 64 of the 270 lines and it reaches the model verbatim; whether a sentence in it can withdraw a ground the tables already prove is unmeasured on this corpus.
The limitWhat a guardrail is not
What exists here is a missing capability plus an instruction. There is genuinely no route through this code that pays, denies or transmits anything — but that is a fact about what was never built, not a control standing guard. No policy engine sits in front of it, no approval chain gates it, no audit trail records it. Kits do not carry those, and bolting them on would turn this into a product rather than something a person can read in an evening.
src/select.py withholds a NAMED SECTION and is not a redaction system. A claim whose Service Correspondence mentions the district manager's goodwill authority will send that mention, because the correspondence is where the answers live and this kit cannot have both. If you point this at real claim packs, that is the first thing to look at.
A paid arm's citations are counted, never constrained. The instruction it is sent forbids naming any provision or record the pack does not print; instructions are things a model may decline to follow, and this one has nothing behind it. The set of identifiers the pack actually contains is already parsed in code and is used only to score, never to reject.
EVERY FLOOR FIGURE HERE HAS 0.05 HOURS BURIED IN IT. That constant was picked by sweeping the same 270 lines the floors are then graded on, and it swings claim-gate by 18.52 points. It is written out beside every number it produced rather than being allowed to pass as a property of the technique.
The 100 per cent on the table half is a BAR AND NOT A RESULT. src/checks.py defines what 'the tables settle' means and the answer key re-derives from it, so that column measures agreement with a definition. The column that measures difficulty is the correspondence half, where free code scores 0 of 22.
THE PAID ARM EXISTS NOW AND IT LOSES THE HEADLINE COLUMN. r001-warranty-claim adjudicates 44.44 pct of 270 lines defensibly against free code's 76.30 pct. It is not a weak arm -- it takes 22 of 22 correspondence grounds and 42 of 42 labour-time cells that free code cannot reach at all, and cuts silent over-denial from 42 lines to 12. It loses on discipline: half the payable lines denied over dealer rounding, 10 of 16 unsettleable lines answered anyway, and a quarter of its citing lines quoting something real that is not a record an analyst can pull.
AND THE CLAIM-LEVEL ACTION METRIC IS BEATEN BY A CONSTANT. 82.22 pct is exactly what the single phrase RETURN_FOR_ADJUSTMENT scores on these 45 claims. It is published rather than dropped, but it should not be read as evidence of anything.
THE 70.74 PCT COUNTERFACTUAL IS A DIAGNOSIS OF 44.44 PCT AND NEVER A REPLACEMENT FOR IT. Computed offline from the committed result file: crediting WP-3.4 gives 56.67 pct, waiving the in-pack non-identifier citations gives 55.93 pct, both together 70.74 pct -- still 5.56 points behind free code. The published figure stays 44.44 pct. The key was not edited, the scorer was not loosened and the run was not fired again; the arithmetic is shown so a reader can see what the gap is made of, not so the gap can be argued down.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 97 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
21 measured by the latest run76 need the model half
Metric
Owner
Role
Why this one
warranty-claim-adjudication-accuracy
The whole adjudication on each claimed line -- disposition, ground, provision, record
alarm
the discriminator against the strongest FREE floor's 76.30 pct, which is the bar and which the paid arm did not clear; the labour-time column split by channel, never blended -- 20 table cells and 22 correspondence cells behave completely differently; the two error directions on their own denominators, never averaged; the tolerance sweep, because the free floors' scores move 18.52 points across it — alarm on the discriminator moving more than one row of its own denominator (0.37 pct on 270 lines) between two runs at the SAME tolerance, the same dataset version and the same arm. A change of tolerance or of arm is a change of system, not a regression.
warranty-claim-verdict
Verdict only, three-way
alarm
the GAP between this and the discriminator, per arm -- it is the citation discipline in one number: 27.78 points on the fast tier and 0.00 on every free floor — alarm on the gap widening on an arm whose verdict accuracy did not move, which means citation quality fell without the judgment changing.
warranty-claim-answer-key-gate
The gate on the answer key itself
alarm
the four seeded defects still convicting by name; the acquittal still passing after the self-test, because a gate that convicts everything convicts nothing — alarm on any violation at all. The baseline is zero and there is no allowance to spend.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
45
different corpus — nothing is comparable
corpus.bytes
455,373
warranty claim review packs edited — the count held, the bytes did not
split.count
270
the claimed lines count moved — a different set was scored
split.size_p50
10,161
the median size of one claimed line moved
split.size_p95
11,094
the 95th-percentile size of one claimed line moved
dataset.rows
270
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (adjudicated_lines 270, answered_trap_lines 42, blind False, claim_support_cells 70, claims 45, claims_answered 45, correspondence_ground_cells 22, dataset_version warranty-claim-2026-08-26-45packs-v2, deniable_lines 130, denial_ground_cells 70, floor_tolerance_hours 0.05, labour_hours_cells 261, labour_time_cells 42, labour_time_prose_cells 22, labour_time_table_cells 20, lines_citing_any_identifier 0, payable_lines 124, policy_provision_cells 70, table_ground_cells 108, unsettleable_claim_cells 16) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Claim adjudication accuracy (the discriminator)
44.44 pct
270 claimed lines across 45 packs
r001-warranty-claim, the fast tier, 120 of 270 -- against 76.30 pct for the strongest free floor (b002-warranty-claim-claimgate) and 48.15 pct for denying every line. A line scores only when the disposition, the ground, the provision and the records are all right and nothing outside the pack's identifier set was cited.
Verdict only -- the same three-way call with the citations dropped
72.22 pct
270 claimed lines
r001-warranty-claim, against 76.30 pct free. THE 27.78-POINT GAP BETWEEN THIS COLUMN AND THE ONE ABOVE IS THE WHOLE FINDING: the fast tier is close to the floor on what to decide and far behind it on what to write down. On the free floor the two columns are the same number, because it only ever states grounds it read out of the pack.
Unpayable lines denied
100.00 pct
130 lines that were not payable
r001-warranty-claim, 130 of 130 -- against 108 of 130 (83.08 pct) free. Never read without the over-denial band below it: denying everything scores 100 pct here too.
Over-denied lines
50.00 pct
124 payable lines
r001-warranty-claim, 62 of 124 -- against 42 of 124 (33.87 pct) free. ALL 62 ARE ONE SHAPE, verified line by line in the committed result: every one is LABOUR_TIME_EXCEEDED, every one cites WP-3.4, and every one is the dealer's own rounding of 0.01 h (23), 0.02 h (13) or 0.03 h (26) above the published operation. 50 land on clean lines, 7 on the coverage-extension trap, 5 on the goodwill trap.
Grounds the pack's own tables settle
100.00 pct
108 table-channel denial grounds
r001-warranty-claim and b002-warranty-claim-claimgate, both 108 of 108. A BAR, NOT A RESULT on the free side: src/checks.py is what defines 'the tables settle it' and the key re-derives from it. That both arms sit on the bar is the point -- nobody should pay for this half.
Grounds only the Service Correspondence settles
100.00 pct
22 correspondence-channel denial grounds
r001-warranty-claim, 22 of 22 -- against 0 of 22 for all three free floors, at every tolerance in the sweep.
Labour time exceeded, both channels
100.00 pct
42 lines whose ground is LABOUR_TIME_EXCEEDED
r001-warranty-claim, 42 of 42 -- against 20 of 42 (47.62 pct) free. STILL NEVER QUOTED WITHOUT ITS SPLIT: on the free side the two halves are 100 pct and 0 pct, and a blended figure lets the strong half hide the dead one.
— labour time visibly above the published operation
100.00 pct
20 table-channel labour lines
r001-warranty-claim and b002-warranty-claim-claimgate, both 20 of 20; the free floor decides it at TOL = 0.05 hours.
— labour time superseded by a combination operation, said only in the correspondence
100.00 pct
22 correspondence-channel labour lines
r001-warranty-claim, 22 of 22 -- against 0 of 22 free. On these lines the hours claimed EQUAL the published time exactly and every printed figure is internally consistent; the published time simply is not the time payable.
Silent over-denial -- lines the pack has already answered, denied anyway
r001-warranty-claim, 12 of 42 -- against 42 of 42 (100 pct) free. On its own denominator, never diluted into the over-denial rate.
Unsettleable lines named rather than answered anyway
37.50 pct
16 lines this pack cannot settle (9 operation codes absent from the guide, 7 lines with no technician story)
r001-warranty-claim, 6 of 16 -- against 16 of 16 free. Of the 10 it did not name, 7 were denied outright and 3 called payable.
Denial ground correct, of the six
98.46 pct
130 stated denials
r001-warranty-claim, 128 of 130 -- against 108 of 108 free. The weakest floor, which decides on the coverage window alone, reaches 82.86 pct, because a wear item, an unapproved fitment and an unauthorised prior repair all look identical to a date subtraction.
Policy provision cited correctly
66.15 pct
130 stated denials
r001-warranty-claim, 86 of 130 -- against 108 of 108 free. 42 of the 44 failures are one disagreement: the arm cited WP-3.4 (time above the published operation is payable only with a hotline authorisation) on every labour cell where the key carries WP-3.1 (labour is payable at the time the guide publishes). See the honesty fields -- the key was not edited and the run was not re-fired.
Required record attached
96.92 pct
130 stated denials
r001-warranty-claim, 126 of 130 -- against 108 of 108 free. One record type per ground under src/checks.SUPPORT_TYPE; data/SOURCES.md records that a real adjudication would want several and this corpus carries one.
Published labour time read correctly -- the use case's second half
91.57 pct
261 lines carrying a published operation time (of 270; 9 operation codes have no guide row)
r001-warranty-claim, 239 of 261, compared to the hundredth of an hour (evals/scoring.HOUR = 0.005) -- against 261 of 261 free, which is a table lookup. Scored SEPARATELY and never folded into the discriminator: reading the hours right and the ground wrong is the failure this kit is about.
Citations that are not an identifier the pack defines, and citations absent from it entirely
26.77 pct not an identifier, 0.00 pct absent
269 lines that cited anything at all
r001-warranty-claim, 72 of 269 and 0 of 269 -- against 0 of 150 and 0 of 150 free, which is structural rather than earned. THE GAP BETWEEN THE TWO RATES IS THE FINDING: nothing was fabricated. All 89 tokens are real strings printed in the pack -- 40 operation codes, 14 special coverage adjustment numbers, 13 hotline case numbers, 10 repair order numbers, 5 goodwill authorisations, 4 instances of the literal heading 'Service Correspondence' and 3 line labels -- quoted into a field that asked for a record an analyst can pull.
Claimed lines left out of the reply
0.00 pct
270 claimed lines
r001-warranty-claim, 0 of 270 across 45 replies, and 0 free. The instruction demands one entry per line in pack order and forbids dropping a line for looking unremarkable; an omitted line scores OMITTED rather than being skipped.
Claim-level action (return for adjustment / request support / recommend pay)
82.22 pct
45 claims
r001-warranty-claim, 37 of 45 -- identical to the free floor and identical to the constant phrase RETURN_FOR_ADJUSTMENT, because 37 of the 45 claims carry at least one denial.
The hour tolerance -- one undeclared constant, swept
57.78 pct to 76.30 pct
270 lines at each of eight tolerances, three floors
b900-warranty-claim-tolerance-sweep. The shipped 0.05 hours is the best value for both strong floors and sits inside the MEASURED band over which the key itself does not move -- [0.031, 0.159] hours, re-derived at five values, changing at 0.16. THE PAID ARM HAS NO SUCH DIAL: all 62 of its over-denials are the rounding this constant exists to absorb, and there is nowhere to set it.
Token volume per call
3,316.00 in / 18,959.91 out per claim
45 calls, one per pack, 0 failures and 0 truncated
r001-warranty-claim; 149,220 input against 853,196 output tokens, of which 818,949 (96.00 pct) is provider-side reasoning re-rolled on every call. Largest single reply 30,562 output tokens, 47.8 per cent of the 64,000-token ceiling in src/draft.py. Free floors draw nothing.
Whole-call latency
p50 153,969 ms, p95 233,504 ms
45 unstreamed calls
r001-warranty-claim, 1,503.4 s of wall clock at 5 concurrent workers, against a 1,800-second socket timeout. READ THIS AS AN UPPER BOUND AND NOT A CLEAN READING: the key was being used by sibling kits at the same time, so queueing on the provider side is inside these numbers and cannot be separated out.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
comply · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-warranty-claim-coveragewindow 2026-08-26
b001-warranty-claim-linesweep 2026-08-26
b002-warranty-claim-claimgate 2026-08-26
absent citation, %
—
0.0
0.0
claim action accuracy, %
73.33
82.22
82.22
claim adjudication accuracy, %
34.81
70.37
76.30
claim line omission, %
0.0
0.0
0.0
claim support completeness, %
0.0
100.0
100.0
claim verdict accuracy, %
60.74
70.37
76.30
correspondence ground caught, %
0.0
0.0
0.0
denial ground accuracy, %
82.86
100.00
100.00
denials caught, %
53.85
83.08
83.08
input tokens, whole run
0
0
0
invented citation, %
—
0.0
0.0
labour time accuracy, %
100.0
100.0
100.0
labour time exceeded caught, %
47.62
47.62
47.62
labour time prose caught, %
0.0
0.0
0.0
labour time table caught, %
100.0
100.0
100.0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
output tokens, whole run
0
0
0
overdeny rate, %
24.19
33.87
33.87
policy provision accuracy, %
0.0
100.0
100.0
silent overdeny, %
71.43
100.00
100.00
table ground caught, %
64.81
100.00
100.00
unsettleable named, %
0.0
0.0
100.0
not a time series No two of these 3 runs measured the same system — they differ on claim_support_cells, denial_ground_cells, floor, lines_citing_any_identifier, policy_provision_cells, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
comply · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c001-warranty-claim-calibration 2026-08-26
r001-warranty-claim 2026-08-26
s001-warranty-claim-notes-blind 2026-08-26
absent citation, %
0.0
0.0
0.0
claim action accuracy, %
50.00
82.22
82.22
claim adjudication accuracy, %
33.33
44.44
45.93
claim line omission, %
0.0
0.0
0.0
claim support completeness, %
100.00
96.92
99.07
claim verdict accuracy, %
83.33
72.22
54.81
correspondence ground caught, %
100.0
100.0
0.0
denial ground accuracy, %
100.00
98.46
99.07
denials caught, %
100.00
100.00
83.08
input tokens, whole run
14343
149220
144027
invented citation, %
45.83
26.77
8.58
labour time accuracy, %
91.67
91.57
100.00
labour time exceeded caught, %
100.00
100.00
47.62
labour time prose caught, %
100.0
100.0
0.0
labour time table caught, %
100.0
100.0
100.0
model latency p50 ms
206068.00
153969.00
142311.00
model latency p95 ms
224055.00
233504.00
212947.00
output tokens, whole run
83836
853196
740049
overdeny rate, %
23.53
50.00
73.39
policy provision accuracy, %
42.86
66.15
80.56
silent overdeny, %
16.67
28.57
100.00
table ground caught, %
100.0
100.0
100.0
unsettleable named, %
—
37.5
62.5
not a time series No two of these 3 runs measured the same system — they differ on adjudicated_lines, answered_trap_lines, blind, claim_support_cells, claims, claims_answered, correspondence_ground_cells, deniable_lines, denial_ground_cells, labour_hours_cells, labour_time_cells, labour_time_prose_cells, labour_time_table_cells, lines_citing_any_identifier, payable_lines, policy_provision_cells, table_ground_cells, unsettleable_claim_cells — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-warranty-claim-stub 2026-08-26
claim action accuracy, %
73.33
claim adjudication accuracy, %
34.81
claim line omission, %
0.0
claim support completeness, %
0.0
claim verdict accuracy, %
60.74
correspondence ground caught, %
0.0
denial ground accuracy, %
82.86
denials caught, %
53.85
input tokens, whole run
108952
labour time accuracy, %
100.0
labour time exceeded caught, %
47.62
labour time prose caught, %
0.0
labour time table caught, %
100.0
model latency p50 ms
0.00
model latency p95 ms
0.00
output tokens, whole run
19989
overdeny rate, %
24.19
policy provision accuracy, %
0.0
silent overdeny, %
71.43
table ground caught, %
64.81
unsettleable named, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 21 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
TOL in src/checks.py
every free-floor score on this page, the table half of the key if it travels far enough, and -- by its absence on the paid side -- half of the fast tier's payable lines.
measured
claim-gate runs 57.78 pct at 0.00 hours to 76.30 pct at 0.05, 18.52 points, non-monotonic and falling away on both sides. The key is IDENTICAL over [0.031, 0.159] hours and changes at 0.16, so no floor is published outside that band whatever the sweep says. AND THE SAME CONSTANT IS WHAT THE PAID ARM HAS NOT GOT: all 62 of r001-warranty-claim's over-denials are 0.01 to 0.03 hours of dealer rounding denied as LABOUR_TIME_EXCEEDED -- 23, 13 and 26 lines at each value, verified one by one in the committed result. A float absorbs it on one side of the comparison and there is nowhere to put it on the other.
MILE_TOL in src/checks.py
the 18 OUT_OF_COVERAGE_MILEAGE lines, and nothing else.
reasoning
It follows from the design: a coverage mileage limit is a number written in a policy, not a measurement with a rounding error, so an hour tolerance run over an odometer reading would be a category error. Fixed at 0 miles and named separately. NOT MEASURED -- there is no sweep behind this edge.
NEVER_SENT in src/select.py
what leaves this machine, and whether the model has been handed the dealer's chargeback exposure, the district manager's goodwill authority and the manufacturer's relationship posture toward the store -- a reason to pay that bears on none of the six grounds.
measured
1 section withheld on all 45 packs; sections_used is ['system', 'pack'] on all four calls of c001-warranty-claim-calibration. What it does NOT catch is the goodwill authority mentioned inside the Service Correspondence, which is sent.
MAX_TOKENS in src/draft.py, and TIMEOUT_S in src/adapters with it
how many replies survive, and whether a truncation defect becomes a transport defect the retry policy pays for twice. They are one setting wearing two names.
measured
the discarded probe drew 30,310 of 32,000 (94.7 pct), so 32,000 was abandoned before any scored arm fired; the re-take at 64,000 drew 27,045 (42.3 pct) with 0 failures on the same four packs. The SAME packs drawing 30,310 then 27,045 is the per-call reasoning re-roll, not headroom.
the six checks and their fixed PRIORITY in src/checks.py
which half of the job is free, which half is left to pay for, and what word two arms use for the same defect.
measured
108 of 108 table grounds free and 0 of 22 correspondence grounds, at every one of the eight tolerances swept, on all three floors.
the technician-story derivation in tools/build_corpus.py
whether a line's records are ABOUT that line, and therefore whether an arm answering INSUFFICIENT_EVIDENCE is wrong or right.
measured
the discarded calibration probe answered INSUFFICIENT_EVIDENCE on 14 of 17 payable lines and said in its own words, line by line, that the record on file described a different component. It was right: stories were drawn from a fixed pool instead of being derived from the line's own operation and cause code. The generator was changed, the dataset moved to -v2, and the probe was re-taken. Its ceiling reading stands; its scores are discarded and published nowhere.
the six-ground vocabulary and the three dispositions
all five artefacts at once. The instruction, the checks, the scorer, the generator and the key each hold their own copy of the vocabulary, so a seventh denial ground is a single change spread across four files.
measured
four seeded defects, all four convicted by name, and the shipped key clean before and after -- red-proven in memory, never on disk.
the Service Correspondence reaching the model verbatim
every correspondence ground, all 42 already-answered lines, and -- under an instruction-shaped sentence -- potentially the grounds the tables already prove.
measured
WHAT THE CHANNEL IS WORTH IS NOW MEASURED AND WHAT AN ATTACK WOULD DO TO IT IS NOT. r001-warranty-claim reads it and takes 22 of 22 correspondence grounds, 22 of 22 combination-operation labour cells and cuts silent over-denial from 42 lines to 12; every free floor scores 0 of 22 because it never opens the section. So the channel is demonstrably load-bearing. Whether a sentence planted in it can withdraw a ground the tables already prove remains UNMEASURED on this kit -- evals/injection.py is written, carries both phrasings verbatim and has never been fired, and no rate is published.
the corpus seed and dataset_version in tools/build_corpus.py
every figure on this page at once. Regenerating rewrites the packs, the key and the stats together.
measured
one regeneration has already invalidated a live run. warranty-claim-2026-08-26-45packs-v2 is the second version, and the probe taken on the first sits in results/discarded/ rather than being spliced forward.
the special coverage adjustment sentence in tools/build_corpus.py
how tidy the reverse trap reads, though not whether it is a trap.
measured
On all 16, the limb the denial actually turns on is reopened, so every one of the 16 is a genuine trap. But both limbs are derived from where the vehicle stands rather than from the coverage term, so 8 of the 16 print the OTHER limb at or below the standard term -- 5 on months, 3 on miles -- which reads as an adjustment shortening the half nobody is arguing about. Found after r001-warranty-claim was already in flight. Regenerating would have thrown away 45 paid calls to tidy a sentence, so the note and its fix sit in the generator and the wart is recorded instead.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
Claim adjudication accuracy (the discriminator)
IT FIRED, AND IT IS THE HEADLINE. Free code beats the paid arm by 31.86 points on the one column that asks for an adjudication a dealer can be shown. The kit leads with that rather than burying it.
Verdict only -- the same three-way call with the citations dropped
IT FIRED. Nearly all of the paid arm's loss is citation discipline, not judgement.
Over-denied lines
IT FIRED, AND IT IS THE SINGLE LARGEST COMPONENT OF THE GAP. Half of every payable line comes back denied over a hundredth of an hour a dealer management system rounded when a technician clocked off.
Grounds only the Service Correspondence settles
—. THIS IS THE COLUMN THE MODEL IS ACTUALLY BOUGHT FOR, and it is the only one where free code scores nothing at all. It is also the column an untested attack surface sits in; see the threat model.
— labour time superseded by a combination operation, said only in the correspondence
—. The named trap, and the fast tier takes all of it. This is where the 31.86-point deficit is NOT.
Silent over-denial -- lines the pack has already answered, denied anyway
IT FIRED, AT A THIRD OF THE FREE RATE. This is the dangerous failure and it is the one column where paying buys a real reduction: 12 denials written to a dealer holding the manufacturer's own document, instead of 42.
Unsettleable lines named rather than answered anyway
IT FIRED HARD. Refusing to answer what the pack does not settle is the single behaviour that separates claim-gate from the weaker floor beneath it, and the fast tier does it on barely a third of the lines. A confident answer where the guide has no row is exactly what a dealer appeal is made of.
Policy provision cited correctly
IT FIRED, AND WHAT IT IS MEASURING IS PARTLY A DISAGREEMENT RATHER THAN AN ERROR. Published unadjusted for exactly that reason.
Citations that are not an identifier the pack defines, and citations absent from it entirely
IT FIRED ON THE FIRST RATE AND EMPHATICALLY NOT ON THE SECOND. A wrong citation and a fabricated one are different failures, and publishing one blended number would have been unfair in one direction and far too kind in the other.
Claim-level action (return for adjustment / request support / recommend pay)
IT FIRED. Three arms and a fixed string all land on the same figure to the decimal. On this corpus the metric measures nothing, and it is printed rather than quietly dropped.
The hour tolerance -- one undeclared constant, swept
IT FIRED. 18.52 points of swing on claim-gate from a number a reader would never see quoted, and the weakest floor's best score in the sweep (39.26 pct at 1.00 hours) is the value at which its labour catch falls from 47.62 pct to 9.52 pct -- it is not improving, it is abandoning the check.
Token volume per call
—. The input side is 14.9 per cent of the tokens moved; trimming the pack would save almost nothing and would cost the correspondence half outright.
Whole-call latency
IT FIRED ON PROVENANCE RATHER THAN VALUE. There is no single-tenant measurement here, no per-stage instrumentation anywhere in the kit, and the free floors finish all 45 packs in 0.1 s recording 0 ms per claim.
NextThe three you would add first
RUN THE FREE FLOORS AND SWEEP THE TOLERANCE BEFORE YOU CALL ANYTHINGclaim-gate settles 76.30 pct of 270 lines defensibly for $0.00 in a tenth of a second, and r001-warranty-claim's 45 paid calls returned 44.44 pct. The sweep that chose the floor's constant is free and offline as well. Had this kit published only the model figure it would have shipped a number 31.86 points below what costs nothing, with nothing on the page to reveal it.
GIVE A PAID ARM THE DEALER-ROUNDING TOLERANCE THE FREE FLOOR ALREADY HASAll 62 of r001's over-denials are the same shape: 0.01 to 0.03 hours above the published operation, denied as LABOUR_TIME_EXCEEDED, citing WP-3.4 every time. Free code absorbs this with one float. The instruction the model is sent never mentions that a dealer management system rounds when a technician clocks off, so it has no constant to set and no reason to infer one. NOT DONE, and it is the largest single component of the 31.86-point deficit.
FIRE THE INJECTION PROBE BEFORE TRUSTING ANY PAID ARM ON THIS CORPUSThe Service Correspondence is untrusted input that cannot be removed -- it is the only evidence for 64 of the 270 lines. evals/injection.py is written, forces the note into every claim in scope, fixes scope before reading any answer and pairs every cell against the arm's own un-injected reply. NOT DONE. It is the largest unmeasured risk on this kit and it costs one call per claim in scope.
ESCALATE EVERY LINE WHERE THE FREE FLOOR AND A PAID ARM DISAGREETheir mistakes point in opposite directions, and both answers are already on screen together for every line. Pure code denies 42 of 42 lines the pack has already resolved and never reaches one of the 22 correspondence grounds; an arm that does read that section is an arm a sentence in that section may be able to steer. Flagging the lines where they part company adds nothing to the bill of a run that is happening anyway.
CONSTRAIN EVERY CITATION TO THE IDENTIFIERS THE PACK ACTUALLY PRINTSsrc/draft.cited_ids() already pulls every provision and record a reply names, and the identifiers the pack genuinely prints are parsed elsewhere in the same codebase. Discarding the rest is a screen between the reply and the score -- no retraining, no prompt surgery. NOT DONE. On r001-warranty-claim it would have addressed 72 of 269 citing lines, and it would have cost nothing in fabrication risk, because 0 of those 269 named anything the pack does not contain.
PUT A RECORD IDENTIFIER ON EVERY CORRESPONDENCE DOCUMENT IN THE CORPUSThe special coverage adjustments, hotline authorisations and goodwill authorisations exist only as sentences in the Service Correspondence and carry no record id anywhere in the pack. An arm that reads one correctly has nothing valid to cite, so reading it right and citing it wrong become the same score. NOT DONE: changing it moves the corpus and invalidates every recorded run against it.
SWEEP THE MILEAGE TOLERANCE ON ITS OWN SCALE, IN MILESsrc/checks.MILE_TOL is 0 and was deliberately kept out of the hour sweep, because running an hour tolerance over an odometer reading is the category error the sweep exists to expose. That is an argument, not a measurement, and 18 of the 270 lines turn on it.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
The scored arm r001-warranty-claim has been taken on all 45 packs at 5 concurrent workers, and the correspondence-blind ablation s001-warranty-claim-notes-blind on the same 45 at 6. Both are published, both are on this page, and both are single runs. There is no schedule here and nothing re-fires on a clock: the three free floors and the tolerance sweep re-score the whole corpus offline for nothing in about a tenth of a second, so the cheap half of this board can be re-taken on any commit, and the paid half is re-taken only when someone decides to spend. A band moving is a prompt to look, never a trigger to re-run. THE CORRESPONDENCE-BLIND ABLATION HAS NOW BEEN SCORED, AND IT LANDS ON THE FREE FLOOR'S NUMBER IN FIVE COLUMNS. s001-warranty-claim-notes-blind ran the same arm over the same 45 packs with the Service Correspondence section truncated out of the prompt and the instruction saying so. Against r001-warranty-claim: correspondence grounds 22 of 22 -> 0 of 22, the correspondence half of the labour-time trap 22 of 22 -> 0 of 22, labour time exceeded overall 100.00 pct -> 47.62 pct, unpayable lines denied 100.00 pct -> 83.08 pct, and silent over-denial 12 of 42 -> 42 of 42. Every one of those five is the strongest free floor's own figure to the decimal. An arm that was pattern-matching on the tables would barely have moved; this one lost exactly the prose cases and kept exactly the table ones, which is the only evidence this kit has that the sentences were read rather than guessed around.
What this cannot tell you
Whether one paid run says anything durable. r001-warranty-claim is 45 calls taken once. Nothing was repeated, and 96.00 pct of its output tokens are provider-side reasoning the provider re-rolls per call -- the component least likely to come back the same. The 31.86-point deficit against free code is large enough that a repeat is unlikely to reverse it; that is a judgement about the size of the gap, not a second measurement.
Whether a sentence in the Service Correspondence can talk an adjudication out of a ground it has already stated. Zero injection calls have been made on this kit. Both phrasings are written out verbatim in evals/injection.py and neither was fired, so there is no rate here -- not a low one.
Repeat variance on anything. Every arm ran exactly once. On the one live probe 95.8 per cent of the output tokens are provider-side reasoning the provider re-rolls per call, which is the component least likely to reproduce, and the same four packs already drew visibly different totals under two ceilings.
Whether the sweep behaves this way anywhere else. Eight constants over 270 lines of a single generated corpus, and that corpus was deliberately built with a gap in it — dealer rounding capped at 0.03 hours, planted over-claims starting at 0.08. The dip and rise the sweep shows belongs to these packs and should not be carried anywhere.
Whether the Dealer Position section changes an answer. It has never been in a request, so there is no pair of runs to compare and no way to say what its absence is worth.
WHAT THE CORRESPONDENCE CHANNEL IS WORTH IS NOW MEASURED AS AN ABLATION, and the entry that said it was not has been replaced. THE CORRESPONDENCE-BLIND ABLATION HAS NOW BEEN SCORED, AND IT LANDS ON THE FREE FLOOR'S NUMBER IN FIVE COLUMNS. s001-warranty-claim-notes-blind ran the same arm over the same 45 packs with the Service Correspondence section truncated out of the prompt and the instruction saying so. Against r001-warranty-claim: correspondence grounds 22 of 22 -> 0 of 22, the correspondence half of the labour-time trap 22 of 22 -> 0 of 22, labour time exceeded overall 100.00 pct -> 47.62 pct, unpayable lines denied 100.00 pct -> 83.08 pct, and silent over-denial 12 of 42 -> 42 of 42. Every one of those five is the strongest free floor's own figure to the decimal. An arm that was pattern-matching on the tables would barely have moved; this one lost exactly the prose cases and kept exactly the table ones, which is the only evidence this kit has that the sentences were read rather than guessed around. AND THE ABLATION EXPOSED A LIMITATION IN THIS KIT'S OWN HEADLINE METRIC, WHICH IS PUBLISHED RATHER THAN QUIETLY LEFT OUT. The blind arm scores HIGHER on the discriminator than the sighted one -- 45.93 pct against 44.44 -- while being strictly worse at the job: its verdict accuracy falls 72.22 -> 54.81 and its over-denials rise 62 -> 91 of 124. The driver was measured, not guessed: blinding removes the bulletin, hotline and goodwill numbers it was quoting into a record-identifier field, so lines carrying a non-identifier citation fall 72 -> 23 and cells citing WP-3.4 where the key carries WP-3.1 fall 42 -> 20. The citation term in the discriminator therefore rewards an arm for having less to say. Both numbers are published and neither is adjusted. What remains unverified is the adjacent question, which is NOT the same one: whether a correspondence ground already stated can be SUPPRESSED by an instruction-shaped note left in place. That is evals/injection.py, it has never been fired against this kit, and security.redteam records it as not attempted.
Whether a keyword floor over the correspondence would close part of the 0-of-22 gap. Deliberately not built, because it would be tuned to the sentences this generator writes and would measure the generator rather than the method.
Latency as anything a reader can plan against. p50 and p95 come from 45 unstreamed calls run 5 at a time against a credential sibling kits were drawing on simultaneously, so provider-side queueing sits inside both figures and cannot be separated out. Treat them as an upper bound. No kit here is instrumented per pipeline stage, so nothing says where inside a call the time went either.
Whether WP-3.1 or WP-3.4 is the right clause for a line claiming more time than the guide publishes. The arm cited WP-3.4 on all 42 labour cells where the key carries WP-3.1, both clauses are printed in the pack, and WP-3.4 is arguably the more precise of the two. THE KEY WAS NOT CHANGED AND THE RUN WAS NOT RE-FIRED -- editing a key after reading an arm's answers is how a benchmark starts agreeing with whatever it just measured. It is recorded as an open disagreement and the 66.15 pct provision figure is published unadjusted.
Whether the reverse traps read as cleanly as they behave. All 16 coverage-extension notes reopen the limb the denial turns on, so the measurement is sound, but 8 of the 16 also print the other limb at or below the standard term, which is an untidy sentence rather than a broken case. Found after the paid run was in flight and recorded rather than regenerated away.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
NO FRAMEWORK, AND THE COST OF THAT CHOICE IS PAID BY THE AUTHOR RATHER THAN THE READER. Nothing is installed: <code>requirements.txt</code> holds no package and states the condition for adding one — some module under <code>src/</code> or <code>evals/</code> has to import it first. No orchestration layer, no vendor SDK, no eval library. The test this is written to pass is whether somebody who has never seen the repository can find the two decisions that actually move the published numbers and change them in an afternoon. On warranty claims those two are a labour-hour tolerance (<code>src/checks.TOL</code>) and the name of the section that is held back (<code>src/select.NEVER_SENT</code>) — a float and a one-element tuple, both of which any framework worth its name would have swallowed into configuration where nobody argues with them.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the model call
src/adapters/__init__.py
an LLM client wrapper or a Runnable
one dict and two functions over urllib. A wrapper would buy streaming, a retry policy and a provider registry. The registry is a dict here; streaming is not wanted, because a non-streamed completion is what makes the socket timeout a measurable setting rather than an invisible one; and the retry policy is the thing a wrapper would most certainly have got wrong. It is deliberately asymmetric -- four attempts on a transient status, exactly ONE on a transport failure -- because a blown socket on an unstreamed 64,000-token generation is not a rate limit, and retrying it four times pays for the whole generation four times. That asymmetry is written out in the file with the reasoning beside it; in a framework it would be two config keys nobody reads.
the withholding of the dealer's own position
src/select.py
a PII/redaction or data-loss-prevention layer
a tuple of one section name and two functions. What a redaction layer would buy is pattern-based scrubbing across the whole pack; what this kit needs is the opposite -- a NAMED section that never leaves, so the question 'what did we send' has a list for an answer instead of a confidence score. The honest limit is stated rather than engineered around: a claim that mentions the goodwill authority inside the Service Correspondence will send it, because the correspondence is where the answers live and the kit cannot have both.
the six structured checks and their hour tolerance
src/checks.py
a rules engine
Six lookups against tables the pack itself prints, two tests for a record that is simply absent, one fixed reading order, and a single float. What an engine gives you is rules somebody can author without touching Python; what this kit is trying to show is that these six are so cheap to write by hand that writing them removes the whole table half of a warranty adjudication from the bill, at nought pounds and a tenth of a second. And the float beside them is worth 18.52 points of published accuracy. Hand that to a configuration file and you have manufactured precisely the thing the tolerance sweep exists to expose: a headline figure with an undeclared parameter inside it.
the reply parse
src/draft.py
a structured-output or schema-validation library
Brace counting, two vocabulary uppercasings, and one coercion of a lone string into a list. Validate this against a schema and it starts rejecting replies that are correct: a disposition typed in lower case, or one record identifier returned without its brackets, is the same adjudication differently punctuated, and marking it wrong publishes a typography complaint as a judgement failure. The thing that IS caught is the thing an analyst would catch — a provision or a record number the pack never printed is counted against the arm, not tidied into shape.
the free floors and the sweep
evals/baseline.py, evals/tolerance.py
a benchmark harness with built-in baselines
three hand-written floors of increasing strength and an eight-value parameter sweep, each a genuine attempt at the job rather than a strawman. A harness would supply a majority-class baseline and stop; on this kit the majority class is printed too (48.15 pct verdict, 82.22 pct claim action) and is not the interesting floor. The interesting floor is the one that refuses to state a ground it cannot support, and no generic harness would have produced it because it encodes the use case's own definition of a defensible adjudication.
the eval run and the result record
evals/run.py, evals/scoring.py
an eval framework
Threads over independent claims, and one JSON file per run. Frameworks in this space compete on dashboards. What has to come out of a warranty run instead is a record another repository can read without permission, and a scorer in which no rate is ever quoted without the population it was taken over. That is why the labour-time figure comes out in three pieces rather than one: blend the 20 lines a table can see with the 22 a table cannot and the strong half hides the fact that the weak half is at zero. No off-the-shelf scorer splits that way, because the split is a fact about combination operations and not about evaluation.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Linear, once per claim pack: data/corpus/WC-n.txt -> src/segment.py (six underlined sections) -> src/select.py (five of them leave) -> src/prompt.py (system plus pack, in send order) -> src/adapters (one call) -> src/draft.py (parse and normalise) -> evals/scoring.py (exact match per line against the key). No branches, no agent loop, no tool calls, no retrieval step and nothing re-entrant. The only fan-out is a thread pool over independent claims, and the three free floors skip the middle four stages entirely -- src/checks.py reads the parsed pack directly.
The other sideWhat a framework costs you
A new provider is a hand-edited entry in PROVIDERS, and one that reports no token usage simply cannot be published here, since the cost and telemetry surfaces both read those counts off the run record. What is bought in return is that cloning this kit installs nothing at all.
There is no configurable retry or backoff policy. It is four attempts with 1s/2s/4s/8s backoff on a transient status and one attempt on a transport failure, written out in the adapter. Changing it means editing Python, which is the intended cost of a decision this consequential.
There is no schema validation on the reply, so a genuinely malformed answer is scored as a miss rather than raised as an error. That is the deliberate trade -- see the reply-parse seam -- and it is only defensible because parse failures are recorded per run rather than assumed away.
There is no persistence layer and no job queue, so a run is a process. A run that dies loses its unrecorded calls, and the daily call ledger over-counts by one on a crash mid-call rather than under-counting -- the direction a spend guard should round.
What we could NOT verify
Whether a structured-output library would raise the parse rate. No arm on this kit has recorded a parse failure, so there is nothing for one to fix -- but the only live run is four calls, which is far too few to say a parse failure cannot happen.
Whether an orchestration framework would change any score. No port was built. The claim here is about what a framework would COST a reader of this code, and it is argued from the seams rather than measured against an alternative implementation.
Whether the asymmetric retry policy is correctly tuned. No transport failure has occurred on any recorded run of this kit, so the ONE-retry rule is reasoning carried over from a sibling kit's measurement, not a measurement taken here.
Whether a keyword floor over the Service Correspondence would close part of the prose gap. It was deliberately not built, because it would be tuned to the sentences this generator happens to write and would measure the generator rather than the method -- so the 0 of 22 is a statement about THESE three floors and not about what free code can do in principle.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-warranty-claim on the fast tier, correspondence-blind, 2026-08-26. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
153,969 ms
p50 153,969 ms, p95 233,504 ms
IT FIRED ON PROVENANCE RATHER THAN VALUE. There is no single-tenant measurement here, no per-stage instrumentation anywhere in the kit, and the free floors finish all 45 packs in 0.1 s recording 0 ms per claim.
Model, p95
233,504 ms
p50 153,969 ms, p95 233,504 ms
IT FIRED ON PROVENANCE RATHER THAN VALUE. There is no single-tenant measurement here, no per-stage instrumentation anywhere in the kit, and the free floors finish all 45 packs in 0.1 s recording 0 ms per claim.
Input tokens
149,220
3,316.00 in / 18,959.91 out per claim
—. The input side is 14.9 per cent of the tokens moved; trimming the pack would save almost nothing and would cost the correspondence half outright.
Output tokens
853,196
3,316.00 in / 18,959.91 out per claim
—. The input side is 14.9 per cent of the tokens moved; trimming the pack would save almost nothing and would cost the correspondence half outright.
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c001-warranty-claim-calibration206,068 ms
r001-warranty-claim153,969 ms
s001-warranty-claim-notes-blind142,311 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-warranty-claim-coveragewindow, b001-warranty-claim-linesweep, b002-warranty-claim-claimgate recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the rulebook is data read whole into the prompt.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-26, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
warranty claim review packs
data/corpus/WC-<n>.txt — 45 files, 455,373 bytes, 10,119 bytes on average, generated from seed 20260826 as dataset warranty-claim-2026-08-26-45packs-v2
five of the six underlined sections go to the provider; Dealer Position never does (src/select.py, NEVER_SENT)
the answer key
data/gold.jsonl — 45 claims, 270 labelled lines. tools/build_corpus.py emits it in the same pass as the packs, and evals/check_labels.py then rebuilds it out of those packs using its own hand-typed case table rather than the generator's
never
the recorded runs
results/eval-*.json, committed — three free floors, an eight-value tolerance sweep, a stub and one live calibration probe. results/discarded/ holds the run that was thrown away and says why
never
the local adjudication UI
src/app.py, bound to a loopback port. It reads claims, prints the assembled prompt, replays the recorded run and computes the free floor with no credential at all; only its draft path calls anything
nothing on the read paths. The draft path sends the assembled pack minus the withheld section
the provider credential
<repo>/.env, this kit's optional .env, or the real environment, in that precedence, read by src/config.py and gitignored from the first commit
to the provider the reader configured, and nowhere else. This repo has never held one, and the arm that scores highest on this corpus needs none
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 79
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The credential is looked for in three places and the most specific one wins: the repository-wide <code>.env</code>, then an optional file belonging to this kit alone, then whatever the real environment holds. Merging happens variable by variable, so a kit-level file naming only MODEL redirects this one adjudicator while continuing to use the shared account. Neither file has ever been tracked — both were ignored from the first commit and no credential has existed in this repository. <code>src/config.sources()</code> answers where a setting came from by naming files and never contents, so that question can be settled without anybody reading a secret out loud. <code>save()</code> creates the file already at 0600, so there is no window at the default umask in which it sits readable. If the provider returns an error, the local UI strips the key and the base URL out of the message before it reaches the browser. A reader is never asked for any of this: the interface loads, browses, replays and computes floors with nothing configured, which on warranty claims is not a crippled mode at all, since the best-scoring arm the kit publishes calls no one.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
rulebook
the whole Warranty Policy And Guides section, sent verbatim inside the pack rather than pre-extracted: coverage terms, cause classification, the labour operation guide and the provision id beside each row. Its integrity is the pack's own -- there is no separately maintained policy file that can go stale against the claims, and a coverage term a special coverage adjustment has already extended is a planted case in the corpus rather than a configuration error in the kit. Nothing is summarised on the way in, because a summary is where the sentence that reopens a coverage gets lost and the measurement quietly becomes a measurement of the summariser.
3,316.00 input tokens per claim on the scored run, and a two-point solve that splits any pack exactly: 1,124.44 fixed tokens plus 0.228411 tokens per pack character. About two thirds of the input bill is the claim pack and one third the instruction. (p001-warranty-claim-prompt-tokens (2 calls at max_tokens=1), r001-warranty-claim (45 calls))
THE INPUT SIDE IS 14.9 PER CENT OF THE TOKENS MOVED. 149,220 input against 853,196 output across 45 calls, and 818,949 of that output -- 96.00 per cent -- is provider-side reasoning re-rolled every call. Trimming the pack to save input would save almost nothing and would cost the correspondence half of the answer, which is the only half the paid arm demonstrably wins.
A claim pack whose policy and labour time guide are SYSTEMS rather than attachments. data/SOURCES.md says what happens: src/claim.py parses no coverage terms and no operations, and every line comes back INSUFFICIENT_EVIDENCE -- the honest failure rather than a hidden one.
model
a single completion per claim pack, aimed at whatever provider and credential the reader already has, leaving from the one outbound function in src/adapters/__init__.py. Before any of that happens, src/checks.py settles every ground the pack's own tables settle — for nothing, in a tenth of a second — so a paid call is only ever asked the part the tables leave open. The daily call ceiling is enforced inside the adapter, ahead of the request being assembled, which means every kit sharing one credential draws on a single ledger.
the strongest free floor adjudicates 76.30 per cent of 270 lines defensibly for $0.00 in 0.1 seconds; the fast tier, given all 45 packs, returns 44.44 per cent. The paid arm wins the two columns free code cannot reach -- 22 of 22 correspondence grounds and 22 of 22 combination-operation labour cells against 0 of 22 on both -- and loses on discipline, denying 62 of 124 payable lines over dealer rounding and answering 10 of 16 lines the pack does not settle. (b002-warranty-claim-claimgate, r001-warranty-claim)
ON THIS CORPUS THE FREE HALF IS NOT A FLOOR, IT IS THE WINNER, AND IT IS ALSO THE HALF NO SENTENCE CAN SWITCH OFF. What it cannot do remains the reason to pay at all: it denies 42 of 42 lines the pack has already answered, because it never reads the answer. The paid arm cuts that to 12 and inherits the opposite ceiling -- it reads the channel, and the channel is untrusted input nobody has yet attacked here.
A denial ground whose record sits in a field the parser has never seen, and an hour tolerance that is not yours. Every free-floor figure here is computed at src/checks.TOL = 0.05 hours.
labels
45 packs and 270 labelled claimed lines, generated in the same pass as the corpus and then re-derived from it by a second reader. Eleven properties are asserted, and the three that matter are the corpus's claims about itself rather than its schema: the table half must re-derive exactly from src/checks.py, a correspondence denial must look PAYABLE to the tables, and a correspondence trap must look DENIED to them.
11 asserted properties, 0 violations on the shipped key; four seeded defects convicted by name under --self-test, seeded in memory and never written to disk. (evals/check_labels.py)
THE TABLE HALF OF THE KEY IS CLOSER TO A DEFINITION THAN A DISCOVERY. src/checks.py is what DEFINES 'the tables settle it', so the free floor's 100 per cent on that half is published as a bar and not as a result. What the assertion buys is that the key and the floor cannot silently drift apart -- not that either is right about a real claim.
Your own claims. Nothing here transfers a label; the six grounds and the three dispositions are this corpus's vocabulary and a real warranty operation has its own.
corpus refresh
the generator, not the data's provenance. tools/build_corpus.py rewrites the packs, the key and the stats from one seed in one pass, so the corpus a reader re-scores is the corpus the published figures were taken on -- and a change to the generator moves the key with it by construction.
one full regeneration has already invalidated a live run: the dataset moved to warranty-claim-2026-08-26-45packs-v2 when every technician story was re-derived from its own line's operation and cause code, and the probe taken on the earlier corpus was discarded rather than published. (results/discarded/README.md; data/corpus-stats.json)
THE VERSION IS A COMPARABILITY GUARD, NOT A LABEL. Every result file records dataset_version, and a run taken on one corpus cannot be differenced against a run taken on another. That is not hypothetical here -- it is what the discarded probe is.
Any published figure whose run record names a different dataset_version, and any comparison across the regeneration.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a DENY with no policy provision and no record identifier attached
the arm has produced an exception list rather than an adjudication. A denial with nothing behind it is not something a warranty analyst can put in front of a dealer -- it comes straight back as an appeal.
read claim_adjudication_accuracy_pct against claim_verdict_accuracy_pct; the gap is exactly the citation and record requirement (evals/scoring.py, and free floor coverage-window which does precisely this -- 34.81 pct adjudication against 60.74 pct verdict, at 0.0 pct provision accuracy on all 70 denials it states)
a provision or record identifier in the reply that is not printed anywhere in the pack
a fabricated citation. It is the first thing the dealer's warranty administrator checks, and one of them discredits the lines around it that were right.
invented_citation_pct against absent_citation_pct; the gap between the two is the arm quoting a real correspondence number into a field that wanted a record identifier (src/draft.py cited_ids() against the pack's own identifier set; every free floor reads its provisions out of the pack via src/checks._prov and records 0 of 150 citing lines)
a line denied OUT_OF_COVERAGE_TIME whose own pack carries a special coverage adjustment, or LABOUR_TIME_EXCEEDED where a hotline case authorised the extra diagnostic time
silent over-denial. The document that reverses the denial is the manufacturer's own, and it is already in the pack the arm just read.
silent_overdeny_pct on its own denominator of 42 answered lines -- never overdeny_rate_pct, which dilutes it across 124 payable lines (b002-warranty-claim-claimgate denies 42 of 42, because free code never opens the section that answers them; r001-warranty-claim denies 12 of the same 42)
a claimed line whose hours EQUAL the published operation time exactly, on a repair order carrying a second operation
the combination-operation trap. Every printed figure on the line is internally consistent and the published time is simply not the time payable; only the correspondence says so.
labour_time_prose_caught_pct on its own 22-line denominator, never the blended labour_time_exceeded_caught_pct (evals/scoring.py's three-way labour split; every free floor scores 0.0 pct on the prose half and 100.0 pct on the table half, blending to 47.62 pct)
a run whose over-denials are all one ground and all one clause
the arm is missing a parameter rather than misjudging the claims. A dealer management system rounds labour when a technician clocks off, and an arm with nowhere to absorb that denies the rounding instead of the claim.
group the over-denied lines by ground and by cited provision before reading the headline rate at all (r001-warranty-claim: 62 of 62 over-denials are LABOUR_TIME_EXCEEDED citing WP-3.4, at 0.01 h on 23 lines, 0.02 h on 13 and 0.03 h on 26)
a free floor's score quoted with no tolerance printed beside it
a score with an undeclared constant inside it. Nothing on the face of the number tells you whether the floor is doing real work or whether 0.05 happened to suit this corpus.
run python3 -m evals.tolerance -- free, offline, three floors by eight values (results/eval-b900-warranty-claim-tolerance-sweep.json; claim-gate runs 57.78 pct to 76.30 pct across the sweep and the weakest floor's best score is the value at which it stops working)
output_tokens_max climbing toward max_tokens on a run record
the cap is close to biting. A truncated reply counts as a failure, remains inside the published population, and is neither spliced out nor quietly re-fired at a higher ceiling.
the failures array and output_tokens_max printed against the cap, on the newest record (evals/run.py; the discarded probe drew 30,310 of 32,000 (94.7 pct) and 32,000 was abandoned before any scored arm fired, while r001-warranty-claim drew a largest reply of 30,562 of 64,000 (47.8 pct) across 45 calls with 0 failures and 0 truncations)
the local UI's per-claim panel listing Dealer Position among the sections that were sent
the withholding seam has been edited or the pack's headings have changed shape. The model would have been handed the dealer's chargeback exposure and the district manager's goodwill authority -- a reason to pay that has nothing to do with the claim.
the UI's what-went / what-stayed panel per claim, and sections_used on the run record (src/select.py NEVER_SENT and src/app.py /api/claim; c001-warranty-claim-calibration records sections_used = ['system', 'pack'] on all four calls)
['Whether either arm reproduces. The scored run r001-warranty-claim covers all 45 packs with 0 failures and 0 truncations, but it was taken exactly once, and 96.00 per cent of its output tokens are provider-side reasoning re-rolled per call. No second paid tier was run either, so nothing here ranks one model against another.', "A LIKE-FOR-LIKE ABLATION IS NO LONGER MISSING -- it was fired and scored after this block was first written, and the entry that said otherwise has been replaced rather than left to read as a gap. THE CORRESPONDENCE-BLIND ABLATION HAS NOW BEEN SCORED, AND IT LANDS ON THE FREE FLOOR'S NUMBER IN FIVE COLUMNS. s001-warranty-claim-notes-blind ran the same arm over the same 45 packs with the Service Correspondence section truncated out of the prompt and the instruction saying so. Against r001-warranty-claim: correspondence grounds 22 of 22 -> 0 of 22, the correspondence half of the labour-time trap 22 of 22 -> 0 of 22, labour time exceeded overall 100.00 pct -> 47.62 pct, unpayable lines denied 100.00 pct -> 83.08 pct, and silent over-denial 12 of 42 -> 42 of 42. Every one of those five is the strongest free floor's own figure to the decimal. An arm that was pattern-matching on the tables would barely have moved; this one lost exactly the prose cases and kept exactly the table ones, which is the only evidence this kit has that the sentences were read rather than guessed around. AND THE ABLATION EXPOSED A LIMITATION IN THIS KIT'S OWN HEADLINE METRIC, WHICH IS PUBLISHED RATHER THAN QUIETLY LEFT OUT. The blind arm scores HIGHER on the discriminator than the sighted one -- 45.93 pct against 44.44 -- while being strictly worse at the job: its verdict accuracy falls 72.22 -> 54.81 and its over-denials rise 62 -> 91 of 124. The driver was measured, not guessed: blinding removes the bulletin, hotline and goodwill numbers it was quoting into a record-identifier field, so lines carrying a non-identifier citation fall 72 -> 23 and cells citing WP-3.4 where the key carries WP-3.1 fall 42 -> 20. The citation term in the discriminator therefore rewards an arm for having less to say. Both numbers are published and neither is adjusted. What is still NOT measured here is repeat variance on either arm: each ran exactly once, and 96.00 pct of r001's output tokens are provider-side reasoning the provider re-rolls per call, so a second run of the identical prompt is not promised to reproduce either figure.", 'Repeat variance. Every arm on this kit ran exactly once, and on the one live probe 95.8 per cent of the output tokens are provider-side reasoning the provider re-rolls per call -- which is the component least likely to reproduce.', 'The tail of the output-token distribution. 45 calls put the largest reply at 30,562 tokens, 47.8 per cent of the ceiling, which is headroom rather than a bound -- the same four packs drew 30,310 under a 32,000 cap and 27,045 under a 64,000 one, which is the per-call re-roll and not a measurement of how heavy a reply can get.', 'Single-tenant latency and the effect of concurrency. The scored run went out 5 workers at a time against a credential sibling kits were using simultaneously, so provider-side queueing is inside p50 and p95 and cannot be separated out. No run was taken at one worker, and none at any other setting.', "Anything about provider-side retention of a claim pack. What leaves this machine is measured; what the provider does with it afterwards is the reader's contract with the provider and was not tested.", 'The mileage tolerance. src/checks.MILE_TOL is fixed at 0 miles and was deliberately kept out of the hour sweep, so its sensitivity is asserted from the argument -- a coverage limit is a number in a policy, not a measurement with a rounding error -- and not from a run.', "Whether any of this survives a real transmitted claim record. src/claim.py parses this corpus's fixed-column layout; against a real dealer management system feed it returns empty tables, and the cause classification check has nothing to key on at all.", "Whether the coverage-extension notes read as cleanly as they measure. On all 16 the limb the denial turns on is genuinely reopened; on 8 of them the other limb prints at or below the standard term, because both limbs derive from the vehicle's position rather than from the coverage term. Found after the paid run was already in flight, and recorded rather than regenerated away -- rebuilding the corpus would have discarded 45 paid calls to tidy a sentence.", 'Which of WP-3.1 and WP-3.4 a warranty operation would actually cite for time claimed above the published operation. The arm chose WP-3.4 on all 42 such cells and the key carries WP-3.1; both are printed in the pack. Nobody with authority over the policy has ruled, the key stands unedited and the run was not taken again.']
The corpus licence, from the Data lens: MIT -- this repository's own licence. The corpus has no third-party component to licence separately: it is written by a generator that ships in the repo, from a seed recorded in data/corpus-stats.json, and every figure in it is invented. Verified by reading every generator input on 2026-08-26. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The whole adjudication on each claimed line -- disposition, ground, provision, record
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineThe whole adjudication on each claimed line -- disposition, ground, provision, record
whether each of the 270 claimed lines got an adjudication you could actually put in front of the dealer appealing it: the disposition the answer key carries, and on a DENY the right ground of the six, the policy provision the pack itself prints for it, every record identifier the ground needs, and nothing cited that is not a provision or record the pack names
$0.00per 1,000 warranty claim review packs
nodata leaves your network
yessame answer every time
MethodHow the test was run
python3 -m evals.run --run-id <id> --yes (or --floor <mode> for a free arm); evals/scoring.py then compares each line against data/gold.jsonl by exact match.
Every grader on these pages scored the same 270 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
Claim
WC-0018
Line
L4
What was repaired
Front drive shaft replacement
What the printed tables settle
2.30 hours claimed against the 2.30 hours the labour operation guide publishes for LO-2519; a covered defect; 14 completed months and 38,632 miles against CV-PWR's 60 months and 60,000 miles; the technician story on file. Every printed figure on the line is internally consistent and src/checks.py returns PAYABLE.
What only the service correspondence says
"Labour operations LO-2404 and LO-2519 are a combination on this model: when both are claimed on one repair order the published time for LO-2519 is superseded by the combination time of 1.20 hours." LO-2404 is claimed on L3 of the same repair order.
Strongest free floor
PAYABLE
-- its ground
none -- it returns PAYABLE, so there is no ground to state
The fast tier
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- the provision it cited
WP-3.4
-- the records it attached
['RC-0018-09']
Answer key
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- its provision
WP-3.1
The named trap, on the channel free code cannot reach. The floor says PAYABLE because every table is satisfied; the arm says DENY on LABOUR_TIME_EXCEEDED because it read the combination note and checked that the partner operation really is on this repair order. It still does not score on the discriminator, because it cited WP-3.4 where the key carries WP-3.1 -- which is the answer-key conviction below, not a reading failure.
Grader
Verdict
Why
The whole adjudication on each claimed line -- disposition, ground, provision, record
incorrect
Gold is DENY / LABOUR_TIME_EXCEEDED / WP-3.1 with record RC-0018-09 required. The fast tier answered DENY / LABOUR_TIME_EXCEEDED and attached the right record -- it read the combination note and checked the partner operation really is on this repair order -- but cited WP-3.4, so this grader scores it a MISS. It is one of the 42 cells where the model convicted the answer key and the key was not changed. The strongest free floor answered PAYABLE and is a miss too, for the opposite reason: every table on the line is satisfied and it cannot see the note.
Verdict only, three-way
correct
Without the citation conditions the fast tier's DENY matches gold exactly, so this line counts as a HIT here and a miss on the discriminator -- the whole 27.78-point gap between the two graders in one row. The free floor's PAYABLE is a miss on both. This is the clearest single line on the kit for what the citation conditions actually cost.
The gate on the answer key itself
not applicable
This grader scores the answer key, not an arm, so it returns no verdict on any line. What it does assert about THIS row is that it is a legitimate correspondence case at all: property 7 requires a correspondence DENY to look PAYABLE to the tables, and src/checks.py does return PAYABLE here -- which is why the line is in the 22-cell correspondence denominator rather than the 108-cell table one.
The formulaWhat it computes
claim_adjudication_accuracy_pct = accurate lines / 270
The analysisWhat it actually did
Model
Result
the fast tier
44.4% claim adjudication accuracy · 14 more measured on this row
the fast tier, correspondence-blind
45.9% claim adjudication accuracy · 14 more measured on this row
strongest free floor
76.3% claim adjudication accuracy · 14 more measured on this row
second free floor
70.4% claim adjudication accuracy · 14 more measured on this row
the assigned free floor
34.8% claim adjudication accuracy · 12 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, derived with the corpus by tools/build_corpus.py and re-derived from the shipped packs by evals/check_labels.py, which asserts eleven properties and is red-proven with --self-test.
These rates are UNKNOWN, on purpose
This grader IS the reference standard, so it has no TPR or TNR of its own: scoring it against itself would be circular and would read as evidence. What can be said about it is said by its own gate, evals/check_labels.py, and by the 42 cells where the model convicted it and the key was not changed.
Watch these
the discriminator against the strongest FREE floor's 76.30 pct, which is the bar and which the paid arm did not clear
the labour-time column split by channel, never blended -- 20 table cells and 22 correspondence cells behave completely differently
the two error directions on their own denominators, never averaged
the tolerance sweep, because the free floors' scores move 18.52 points across it
Alarm on
the discriminator moving more than one row of its own denominator (0.37 pct on 270 lines) between two runs at the SAME tolerance, the same dataset version and the same arm. A change of tolerance or of arm is a change of system, not a regression.
How tight can the band be? src/checks.TOL = 0.05 hours, and it was chosen from b900-warranty-claim-tolerance-sweep rather than typed: claim-gate reads 57.78 pct at 0.00 h, 65.56 at 0.02, 76.30 at the shipped 0.05, 75.93 at 0.25 and 74.81 at 1.00. It is published inside the band [0.031, 0.159] hours over which the ANSWER KEY is identical -- measured by re-deriving all 270 labels at 0.031, 0.05, 0.10, 0.15 and 0.159, and it changes at 0.16 -- so the floor cannot move the labels in its own favour.
Cadence: On demand and free. All four arms re-score offline in about a tenth of a second, so this grader can run on every commit; the paid arms are re-taken only when someone decides to spend.
The decisionWhen to reach for it
Use it
Always. Every published figure in this lens rests on it.
Do not use it
It cannot tell a reading that is defensible-but-different from a wrong one, and on this kit it got that wrong 42 times in one shape: every labour-time denial cites WP-3.4 where the key carries WP-3.1, and WP-3.4 is arguably the better clause. All 42 are scored as misses. It also penalises the 72 lines that cited a real bulletin, hotline case or operation code into a record-identifier field exactly as harshly as a fabrication, although 0 of them are absent from the pack. And this kit's own ablation showed the metric can be IMPROVED by giving the arm less to cite: the correspondence-blind arm scores 45.93 against the sighted 44.44 while being strictly worse at the job.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineVerdict only, three-way
the same three-way call WITHOUT the citation and record requirements -- what a rules engine can score. The gap between this and the row above is the part of the work a coverage-window check does not do, in numbers: 72.22 against 44.44 on the paid arm, and 76.30 against 76.30 on the free floor, which cites correctly by construction
$0.00per 1,000 warranty claim review packs
nodata leaves your network
yessame answer every time
MethodHow the test was run
The same run and the same scorer; claim_verdict_accuracy_pct is computed beside the discriminator on the identical 270 lines, without the citation and record conditions.
Every grader on these pages scored the same 270 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
Claim
WC-0018
Line
L4
What was repaired
Front drive shaft replacement
What the printed tables settle
2.30 hours claimed against the 2.30 hours the labour operation guide publishes for LO-2519; a covered defect; 14 completed months and 38,632 miles against CV-PWR's 60 months and 60,000 miles; the technician story on file. Every printed figure on the line is internally consistent and src/checks.py returns PAYABLE.
What only the service correspondence says
"Labour operations LO-2404 and LO-2519 are a combination on this model: when both are claimed on one repair order the published time for LO-2519 is superseded by the combination time of 1.20 hours." LO-2404 is claimed on L3 of the same repair order.
Strongest free floor
PAYABLE
-- its ground
none -- it returns PAYABLE, so there is no ground to state
The fast tier
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- the provision it cited
WP-3.4
-- the records it attached
['RC-0018-09']
Answer key
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- its provision
WP-3.1
The named trap, on the channel free code cannot reach. The floor says PAYABLE because every table is satisfied; the arm says DENY on LABOUR_TIME_EXCEEDED because it read the combination note and checked that the partner operation really is on this repair order. It still does not score on the discriminator, because it cited WP-3.4 where the key carries WP-3.1 -- which is the answer-key conviction below, not a reading failure.
Grader
Verdict
Why
The whole adjudication on each claimed line -- disposition, ground, provision, record
incorrect
Gold is DENY / LABOUR_TIME_EXCEEDED / WP-3.1 with record RC-0018-09 required. The fast tier answered DENY / LABOUR_TIME_EXCEEDED and attached the right record -- it read the combination note and checked the partner operation really is on this repair order -- but cited WP-3.4, so this grader scores it a MISS. It is one of the 42 cells where the model convicted the answer key and the key was not changed. The strongest free floor answered PAYABLE and is a miss too, for the opposite reason: every table on the line is satisfied and it cannot see the note.
Verdict only, three-way
correct
Without the citation conditions the fast tier's DENY matches gold exactly, so this line counts as a HIT here and a miss on the discriminator -- the whole 27.78-point gap between the two graders in one row. The free floor's PAYABLE is a miss on both. This is the clearest single line on the kit for what the citation conditions actually cost.
The gate on the answer key itself
not applicable
This grader scores the answer key, not an arm, so it returns no verdict on any line. What it does assert about THIS row is that it is a legitimate correspondence case at all: property 7 requires a correspondence DENY to look PAYABLE to the tables, and src/checks.py does return PAYABLE here -- which is why the line is in the 22-cell correspondence denominator rather than the 108-cell table one.
44.4% claim adjudication accuracy · 14 more measured on this row
the fast tier, correspondence-blind
45.9% claim adjudication accuracy · 14 more measured on this row
strongest free floor
76.3% claim adjudication accuracy · 14 more measured on this row
second free floor
70.4% claim adjudication accuracy · 14 more measured on this row
the assigned free floor
34.8% claim adjudication accuracy · 12 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl; only the conditions differ.
These rates are UNKNOWN, on purpose
It is not the reference standard and it is not scored against one: it is a strictly weaker view of the SAME comparison, so there is nothing independent to measure it against. Its value is the difference it exposes, not its own level.
Watch these
the GAP between this and the discriminator, per arm -- it is the citation discipline in one number: 27.78 points on the fast tier and 0.00 on every free floor
Alarm on
the gap widening on an arm whose verdict accuracy did not move, which means citation quality fell without the judgment changing.
How tight can the band be? No threshold. It is an exact three-way string match with no tolerance in it.
Cadence: On demand and free. All four arms re-score offline in about a tenth of a second, so this grader can run on every commit; the paid arms are re-taken only when someone decides to spend.
The decisionWhen to reach for it
Use it
Whenever you want to know how much of the gap is judgment and how much is citation discipline. On the paid arm it is 72.22 against the discriminator's 44.44 -- 27.78 points of the gap are the citation and provision conditions alone.
Do not use it
Never on its own, and never as the headline. It scores a verdict a dealer can appeal with no provision behind it as a hit, which is precisely the adjudication a warranty analyst cannot send. It also flatters the paid arm relative to the free floor, whose two figures are identical at 76.30 because it cites correctly by construction.
Decide a car dealer's warranty claim, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineThe gate on the answer key itself
whether the key everything else is scored against is internally sound: eleven properties, red-proven with four seeded defects that must each be convicted by name and an acquittal that must still pass afterwards
$0.00per 1,000 warranty claim review packs
nodata leaves your network
yessame answer every time
MethodHow the test was run
python3 -m evals.check_labels (and --self-test to red-proof it)
Every grader on these pages scored the same 270 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
Claim
WC-0018
Line
L4
What was repaired
Front drive shaft replacement
What the printed tables settle
2.30 hours claimed against the 2.30 hours the labour operation guide publishes for LO-2519; a covered defect; 14 completed months and 38,632 miles against CV-PWR's 60 months and 60,000 miles; the technician story on file. Every printed figure on the line is internally consistent and src/checks.py returns PAYABLE.
What only the service correspondence says
"Labour operations LO-2404 and LO-2519 are a combination on this model: when both are claimed on one repair order the published time for LO-2519 is superseded by the combination time of 1.20 hours." LO-2404 is claimed on L3 of the same repair order.
Strongest free floor
PAYABLE
-- its ground
none -- it returns PAYABLE, so there is no ground to state
The fast tier
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- the provision it cited
WP-3.4
-- the records it attached
['RC-0018-09']
Answer key
DENY
-- its ground
LABOUR_TIME_EXCEEDED
-- its provision
WP-3.1
The named trap, on the channel free code cannot reach. The floor says PAYABLE because every table is satisfied; the arm says DENY on LABOUR_TIME_EXCEEDED because it read the combination note and checked that the partner operation really is on this repair order. It still does not score on the discriminator, because it cited WP-3.4 where the key carries WP-3.1 -- which is the answer-key conviction below, not a reading failure.
Grader
Verdict
Why
The whole adjudication on each claimed line -- disposition, ground, provision, record
incorrect
Gold is DENY / LABOUR_TIME_EXCEEDED / WP-3.1 with record RC-0018-09 required. The fast tier answered DENY / LABOUR_TIME_EXCEEDED and attached the right record -- it read the combination note and checked the partner operation really is on this repair order -- but cited WP-3.4, so this grader scores it a MISS. It is one of the 42 cells where the model convicted the answer key and the key was not changed. The strongest free floor answered PAYABLE and is a miss too, for the opposite reason: every table on the line is satisfied and it cannot see the note.
Verdict only, three-way
correct
Without the citation conditions the fast tier's DENY matches gold exactly, so this line counts as a HIT here and a miss on the discriminator -- the whole 27.78-point gap between the two graders in one row. The free floor's PAYABLE is a miss on both. This is the clearest single line on the kit for what the citation conditions actually cost.
The gate on the answer key itself
not applicable
This grader scores the answer key, not an arm, so it returns no verdict on any line. What it does assert about THIS row is that it is a legitimate correspondence case at all: property 7 requires a correspondence DENY to look PAYABLE to the tables, and src/checks.py does return PAYABLE here -- which is why the line is in the 22-cell correspondence denominator rather than the 108-cell table one.
The formulaWhat it computes
evals/check_labels.py --self-test -> 0 violations on the shipped key
The analysisWhat it actually did
Model
Result
the shipped answer key
no headline metric on this row — it records violations 0 · properties asserted 11 · seeded defects convicted 4 · acquittal after self test clean
In operationWhat to monitor
Reference standard: src/checks.py plus the shipped corpus -- the key is re-derived from the rendered packs, never from the generator's variables.
These rates are UNKNOWN, on purpose
It grades the key rather than an arm, so it has no TPR or TNR against an arm's truth. Its own red-proof is the only evidence about it, and that is reported as counts of convictions rather than as a rate.
Watch these
the four seeded defects still convicting by name
the acquittal still passing after the self-test, because a gate that convicts everything convicts nothing
Alarm on
any violation at all. The baseline is zero and there is no allowance to spend.
How tight can the band be? No threshold. Eleven boolean properties, all or nothing.
Cadence: On demand and free. All four arms re-score offline in about a tenth of a second, so this grader can run on every commit; the paid arms are re-taken only when someone decides to spend.
The decisionWhen to reach for it
Use it
Before believing any arm. Every percentage on this page is a comparison against the key, and a mislabelled line produces a confident number no gate on the harness can see.
Do not use it
It cannot tell you the key is RIGHT, only that it is internally consistent and that it agrees with src/checks.py. Property one is closer to a definition than a discovery -- src/checks.py is what DEFINES the table half -- which is exactly why the free floor's 100 pct on that half is published as a bar and not as a result. And it did not catch the WP-3.1/WP-3.4 provision choice, because both are printed in every pack and both pass every property it asserts.
A living map of modern AI — kept current every morning