⚠︎ This kit has not been graded
0 completed runs, none scored.
Everything below is coverage, latency and cost. No figure on any of these pages says whether a brief is any good.
The business caseThe problem this solves
A member turns up with two coverages on file and two records of each of them, written by different parties and never reconciled. Somebody has to decide which payer is billed first and name every field the two records disagree about, before the claim goes out. Get the order wrong and the claim comes back denied for other coverage and is reworked; miss the disagreement and the same denial arrives from the payer's side instead. The two questions are usually asked by different people, out of different screens, and neither answer is written down anywhere a later reader can check. Opening two records of the same coverage side by side, deciding field by field whether a hyphen, a corporate suffix or a leading zero is a difference or a spelling, then applying the plan's order-of-benefits schedule from the top and stopping at the first rule that matches — while remembering that the birthday rule is a month-and-day test and that a coverage which terminated before the date of service takes no position at all.
Audience
A benefits coordinator or patient-access lead working a pre-billing worklist — the person who reads the row and either bills it or picks up the phone. Not a claims examiner and not an adjudicator: nothing here prices, holds or pays anything. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual coordination-of-benefits review packs
The corpus is 60 coordination-of-benefits review packs, 0.13 MB (txt 60). ⚠︎ EVERY PACK IS INVENTED AND THERE IS NO REAL PATIENT, MEMBER, SUBSCRIBER, PAYER, PLAN SPONSOR OR CLAIM DATA IN THIS CORPUS ANYWHERE. No person is named — a member is a number — and no payer, employer, plan sponsor, clinic or programme that exists is referred to. Every pack is generated in-process by tools/build_corpus.py from the fixed seed 20260908, and that script FAILS THE BUILD on an email address, a social-security-shaped identifier, a telephone number or a person's name behind a title, so the claim is asserted at build time rather than made in a README. COB-2026 is an invented plan sponsor's CONTRACT schedule: not a law, not a regulation, not a model rule and not any payer's actual contract. It is generated because the answer has to be DERIVABLE rather than typed — each pack is built as a structure (which coverage is in force, who the subscriber is, which subscriber is the active employee, whose birthday falls earlier in the calendar year, which field the front desk keyed wrong, and what shape the intake record is in) and the key is COB-2026 applied to that same structure. And it is generated because it could not be collected: every real pack of this kind is two records about one person's insurance.
The corpus
- The 60 coordination-of-benefits review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md, which also attacks the generator rather than describing it: the notes are templated enough that a regex reading them reaches the key exactly (and evals/check_labels.py is that program, agreeing on 120 of 120), one form variant per pack makes the equivalence rule easier than reality, two families are one line of code away, every pack has exactly two coverages, and the payer pool is fifteen names.
Swap this folder for your own material and the kit is pointed at your coordination-of-benefits review packs. That is the whole change — there is no database to migrate.
==============================================================================================================
COORDINATION OF BENEFITS REVIEW PACK COB-0001
Plan sponsor: Adair Regional Care Network (invented)
Encounter: 2026-05-28 telehealth visit Member: MBR-408000
Schedule: COB-2026
==============================================================================================================
REGISTRATION INTAKE RECORD
What the front desk keyed at scheduling, from the cards the member presented. This is
the only record in this pack that carries the RELATION, the SUBSCRIBER DOB and the
STATUS of the subscriber's employment.
DOC CVG PAYER MEMBER ID GROUP PLAN TYPE RELATION SUBSCRIBER DOB STATUS EFFECTIVE TERM
REG-330000 A Fenwick Mutual, Inc. FM4400001 20000 COMMERCIAL CHILD 1971-01-06 ACTIVE 2019-01-01 -
REG-330001 B Cascade Mutual Health, Inc. CMH4400000 10000 COMMERCIAL CHILD 1995-08-12 ACTIVE 2020-06-01 -
PAYER ELIGIBILITY RESPONSE
What each payer returned for this member. The payer holds the policy, so its PLAN
TYPE and its dates are what the order of benefits is applied to.
DOC CVG PAYER MEMBER ID GROUP PLAN TYPE EFFECTIVE TERM
ELG-770000 A Fenwick Mutual FM4400001 20000 COMMERCIAL 2019-01-01 2026-04-28
ELG-770001 B Cascade Mutual Health CMH4400000 10000 COMMERCIAL 2020-06-01 -
PRIOR COORDINATION LOG
What has already happened on this member's file. It is history, not a rule.Abridged — the file continues.
The outcomeWhat a good result looks like
One pack in, two coordinator's rows out: for each coverage the FIRST rule of COB-2026 it matches, the position that rule gives it, the set of fields on which its two records disagree, and the document rows the finding turns on copied verbatim out of the pack. Every one of the four is graded on every coverage.
And when it cannot
And what it does when it cannot. 60 of 60 calls returned a parseable reply, nothing was truncated and no call errored, so this kit has no measured failure-to-answer rate at all — the failures it does have are wrong answers rather than absent ones, and they are counted one by one. A reply that will not parse is returned as a failure carrying its own token counts and stays inside every published denominator, because it was billed. A coverage a reply did not answer is scored zero and the station answers None for it rather than reaching the terminal rule, so a silence is never scored as a coordinated coverage.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your packs are clean columnar exports and a front-desk note never corrects a value or a status — the free column floor alone — evals/baseline.py,
--floor rules
It settles 42 of 60 packs completely for $0.00, gets every position right on the 42 packs with no reading shape, and misses ZERO of the 31 real disagreements. An exact normalisation and an ordered rule table are things code does not get wrong. - Your intake records carry front-desk notes that re-state a subscriber's employment status, relationship or dates — route those packs to the model and run the free parser on the rest
12 of 12 positions on exactly those packs against the parser's 0 of 12. That is the only slice on this corpus where a reader beats a column, and it is 10 pct of the packs. - You want the discrepancy list itself — the free column floor, and read data/SOURCES.md first
The paid arm reported 61 disagreements that are not there, against 31 that are. The comparison is an exact normalisation, and asking a model to decide whether two spellings are one value is asking it to do badly what code does exactly.
And where nothing here is good enough:
- Three or more coverages on one member's file — neither. This kit is out of scope.
Every pack here has exactly two coverages, the position vocabulary has no TERTIARY and the rule table is pairwise. A three-coverage case would be answered rather than refused, and answered wrongly.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt and data/coverages.json with your own packs and register in the same shapes. You will have no answer key, so nothing can be scored and every accuracy figure on this page becomes inapplicable to your data rather than transferable to it. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for any of it. It is most of the job and none of the value. That is the case against the best-fitting scenario (“Your packs are clean columnar exports and a front-desk note never corrects a value or a status”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | THREE OR MORE COVERAGES ON ONE FILE. Every pack here has exactly two. 4 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the cited row is the row a benefits coordinator would have quoted. The key names the documents COB-2026's cite rule names — both rows of a coverage that disagrees, the eligibility row alone for one that is not in force — and a coordinator might also want the correction note beside an intake row. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-08 — r001-cob-discrepancy. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone, run python3 tools/build_corpus.py, open the board. No key, no install, no index and no network — the 60 packs, the answer key, the coverage register and all five committed runs are in the repository, and every module under src/ and evals/ is standard library.








