The business caseThe problem this solves
A commitment promises a count, in a named placement, inside a window, on written terms. A delivery record then reports what actually ran — usually written by whoever owed the delivery. Somebody reads the record and the schedule side by side and decides whether anything is owed. The question is never IS THERE A GAP. It is four narrower ones, and only the first is a reading problem: which promise is this record even about, does the delivery count against that promise at all, and does the gap survive the three things the agreement itself offers against it — the surplus it lets you carry from another window, the cause it excuses, and the tolerance it wrote for itself. The desk's manual reconciliation of one delivery record against the schedule — the join to the right commitment, the qualifying test, the gross shortfall, the carry-forward, the pro-rated excuse, the tolerance and the remedy ladder — before a person confirms the entitlement record.
Audience
The commercial desk that confirms the entitlement record, and the lead who answers for a compensation given away on a line that met its terms — or for an entitlement that aged out of every window that could have remedied it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual delivery records, each reporting against one commitment
The corpus is 60 delivery records, each reporting against one commitment, 0.03 MB (csv_row 15 · delivery_report 15 · desk_note 15 · platform_export 15). A defect mix a commercial desk would recognise and no public dataset can provide, because five things in it look like shortfalls and are not: a gap already covered by surplus in another window the terms permit you to carry; a gap the terms excuse, pro-rated across the days lost; a gap inside the tolerance the agreement wrote for itself; delivery that reconciles perfectly and ran in a placement the schedule never bought; and a report stated in a unit the commitment does not promise. And two that look settled and are not: two commitments on one line whose windows overlap, where the record's own wording separates them — and one where it does not. The free difference column reads every one of these correctly and answers the wrong question with it: it over-claims on 22 of the 34 records that owe nothing and under-claims on 5 of the 26 that do.
The corpus
- The 60 delivery records, each reporting against one commitmentgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your delivery records, each reporting against one commitment. That is the whole change — there is no database to migrate.
From: delivery-ops@harrowgate-media.example
To: commercial-desk@harrowgate-media.example
Subject: [DELIVERY] MG-0001 - end of flight summary, LN-5255
END-OF-FLIGHT DELIVERY SUMMARY
Record id: MG-0001
Advertiser: Vireo Financial
Campaign: CLEARANCE
Schedule line: LN-5255
Flight: 2026-07-01 to 2026-07-31
Placement delivered: SPORTS PRE-ROLL
Delivered: 1,200,000 impressions
Shortfall cause: none given
Passed to the commercial desk for an entitlement decision. This summary reports what was
delivered; what was promised is on the schedule.
The outcomeWhat a good result looks like
A worked entitlement record a person confirms instead of a difference column a person argues with: OWED, NOT_OWED or DISPUTED; which commitment it is about; the shortfall in whole units, which is what SURVIVES rather than the gross difference; the money that shortfall is worth at the line's own rate; the rung of the remedy ladder whose condition holds; the rule that fired; and one sentence saying what closed the gap or what did not.
And when it cannot
TWO DIRECTIONS AND THEY COST DIFFERENTLY. A FALSE OWED is compensation given away — inventory placed for nothing, or a credit raised against a line that met its terms. A FALSE CLEAR is an entitlement nobody ever claims. Measured over 60 records: the model over-claims $14,400.00 on one record and under-claims $800.00 on one; the honest free arm over-claims $29,638.48 on seven and under-claims nothing at all. ⚠︎ AND $14,400.00 AGAINST $800.00 IS ONE RECORD AGAINST ONE RECORD, NOT EIGHTEEN AGAINST ONE: MG-0014 is priced per spot and MG-0037 per thousand impressions, so a 12-unit error costs eighteen times what a 25,000-unit error does. Any figure that totals dollars across the two bases is measuring the unit mix as much as the arm.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Delivery reporting that already arrives structured — a system that states the line, the flight, the delivered count, the unit, the placement and a coded cause per row — free code, and no model at all
Everything downstream of the reading is ALREADY FREE and measured so: with the same recheck in the path, free code scores 51 of 60 records, gets the release arithmetic right on 6 of 6 inside_tolerance, 5 of 5 cross_period_offset, 5 of 5 placement_missed, 4 of 4 just_outside_tolerance, 3 of 3 unit_mismatch and 3 of 3 offset_not_permitted, claims 26 of 26 entitlements and never fails to claim one. The join, the offset, the pro-rated excuse, the tolerance, the ten rules and the remedy ladder are src/rules.py's work and cost nothing. - Delivery reports that arrive as prose and desk notes, where the reason for a shortfall is a sentence rather than a field — the model, rechecked
This is the only thing the money buys here, and it is worth naming exactly:cause_codeis read on 60 of 60 records against free code's 50, and every one of the eight records the model uniquely wins is a cause record — 4 excused_late_materials (5 of 5 against 1), 2 partly_excused (4 of 4 against 2), 2 cause_not_covered (2 of 2 against 0). That is worth $14,438.48 less money error and six fewer compensations given away. Fourteen of the corpus's seventeen case types are a TIE.
And where nothing here is good enough:
- A book where the same line is re-committed window by window and two commitments regularly overlap one flight — neither, yet — measure it first
That is the axis this kit measures WORST and the only one where the paid arm LOSES: 2 of 4 against free code's 3 of 4, on a denominator of four. Both of the paid arm's remaining errors live there, in opposite directions, and SOURCES.md says plainly that a 50 per cent two-candidate score is two coin flips reported as a rate. The field that separates the two candidates — the campaign name — is printed, load-bearing and NOT GRADED.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own delivery records as .txt into data/corpus/, add one row per record to data/records.json (id, format, assessed_at, source, case, hard), and replace data/schedule.json with your own agreements, terms sets, commitments and recorded positions — the tolerances, the carry-forward permissions and the excused causes are READ, never hard-coded. ⚠︎ A DELIVERY RECORD NAMES AN ADVERTISER, QUOTES A NEGOTIATED RATE AND ASSERTS AN ENTITLEMENT, AND THE WHOLE RECORD PLUS A SLICE OF YOUR OWN SCHEDULE REACHES YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per record for arithmetic a script already does correctly for $0.00. That is the case against the best-fitting scenario (“Delivery reporting that already arrives structured — a system that states the line, the flight, the delivered count, the unit, the placement and a coded cause per row”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL DELIVERY RECORD. Four consistent generated formats from small phrase pools, at a mean of 465 bytes, one record shape per format, sent verbatim. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER ANY OF THIS SURVIVES A SECOND RUN. One scored run, one model, one pass, with provider-side reasoning left at the tier's default and re-rolled per call. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-makegood-candidate. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on a copy of the kit with no .env and no API key, on this machine: python3 -m tools.build_corpus rebuilt the 60 delivery records, the schedule and the key in 0.03 s and every generated file came back byte-identical to the committed ones — repeated under PYTHONHASHSEED 1, 97531, 0 and 12345, with diff -rq empty every time; python3 -m evals.check_labels re-graded that key with 672 independent assertions in 0.03 s and printed KEY CLEAN; python3 -m evals.run --floor rules scored all 60 records in 0.08 s and reproduced the committed floor result's score block exactly (the only per-record difference is the wall-clock ts stamp). Clone to first result with no key: under a second.



