Check what a prior-authorization packet actually proves
A submitted packet says every requirement has a document filed against it, but that column doesn't say if the document is the right kind or still inside the policy's window. This app checks each one and says which requirements actually hold up.
PresenterOpens the private repo. Visible to admins only.
For the plan's review teamHealthcare
Why it matters
Today's manual process, and the same job with the app
A health plan's review team, checking a submitted prior-authorization packet before the clinical determination.
✕Today's manual process
1Open the packet and check the criteria table for which requirements show a filed document.
2Read each document against the policy's kind, window, and minimum, one requirement at a time.
3Check the correspondence for anything that changes whether a requirement still holds.
4A missed gap reaches clinical review and slows the case down, or a valid one gets sent back by mistake.
Every packet read against the table alone
✓With the app
1The packet is read and every requirement is matched to its filed document automatically.
2Each document is checked against the policy's kind, window, and minimum, right there.
3The correspondence is read too so a fact recorded nowhere else still counts.
4Every requirement gets a call met, stale, or unevidenced, with the reason and the clause behind it.
Every packet read against the correspondence too
See it work
One real case, read by the app, step by step
One packet, six requirements: one filed note turns out too old, and a later note withdraws another's evidence.
Check what a prior-authorization packet actually provesReference appBuilt to be shaped to your process
5
1What came in Six requirements, one flag written nowhere but the correspondence.
2The document on file SD-0014-05 is filed and in window, but a later note withdraws it.
3A routine check The tobacco-cessation note is outside the six-month lookback, so this one is returned.
4The record that changed A later addendum withdraws the imaging report, so nothing confirms this one now.
5Not just stale This note is stale too, but an enhanced-review flag blocks the usual equivalent-document waiver.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check what a prior-authorization packet actually proves
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A prior-authorisation packet arrives with a criteria table and a folder of documents. The table says which criteria have a document filed against them, perfectly, and that column is not the answer: a therapy note eleven months outside a twelve-month lookback is filed, a progress note asserting the imaging was performed is filed, a 148-day programme against a 182-day minimum is filed. All three rows read as complete. Somebody has to say which criteria are actually evidenced, by the kind of record the policy names, inside the window it allows -- and then, separately, whether a gap is one the policy already forgives. The manual first-pass read of a submitted packet against the plan's criterion set, before clinical review.
Audience
The plan's review team, before the clinical determination -- and the requesting practice, which answers it criterion by criterion. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual prior-authorisation packets (one request each, 5-6 criteria)
The corpus is 40 prior-authorisation packets (one request each, 5-6 criteria), 0.22 MB (txt 40). A defect mix that a real payer's queue would not give you on demand. 84 criteria settled by the printed tables, 34 only by the clinical correspondence, and 56 traps that look exactly like gaps on the table and are not -- including the MP-6.1 pair, which is identical on the criteria table and differs only in whether the service is enhanced-listed.
The corpus
The 40 prior-authorisation packets (one request each, 5-6 criteria)generated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your prior-authorisation packets (one request each, 5-6 criteria). That is the whole change — there is no database to migrate.
One prior-authorisation packets (one request each, 5-6 criteria), as the model receives itPA-0001.txt · 1 of 40
Prior Authorisation Evidence Check
--------------------
SYNTHETIC RECORD -- invented for an open kit. No real member, patient,
payer, plan, clinician, facility, record or person appears in it.
Packet PA-0001
Request RQ-2026-0101
Plan Northhaven Care (PL-NHV)
Member MB-0001
Requesting provider PR-0001
Service requested Lumbar fusion, single level (SVC-LFUS1)
Submitted 2026-06-25
Reviewed as of 2026-06-29
Enhanced review not recorded on this header
Medical Policy And Evidence Terms
--------------------
Evidence policy
MP-1.1 A criterion is met only where a document filed with this packet
states the fact the criterion requires.
MP-1.2 The filed document must be of the KIND the criterion names. A
statement in a progress note that a study was performed, or that
it supported the request, is not that study.
MP-2.1 Every criterion carries a lookback window, measured back from the
submission date. Evidence dated outside that window does not meet
the criterion, however clearly it states the fact. Where a record
states both a performed date and a report date, the PERFORMED date
is the one this clause runs on.
MP-2.2 Where a criterion carries a minimum duration, the filed evidence
must show that duration completed on or before the submission date.
MP-3.1 A later document that corrects or withdraws an earlier one governs
from its own date, however the criteria table still reads.
MP-4.1 Where a contraindication to a required trial is documented, the
Abridged — the file continues.
The outcomeWhat a good result looks like
An evidence check a reviewer reads first: every criterion with a disposition, the clause it turns on, the documents it rests on, and whether it sends the packet back.
And when it cannot
A check that returns every packet with anything imperfect about it. The practice stops reading, and the review queue is back to a column of missing cells.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Your criteria table is clean and every document is filed beside it, correctly dated — the free floor (evals/baseline.py, evidence-gate) It takes 100.0 pct of the grounds the printed tables prove, cites the clause, attaches the document and applies MP-6.1's first limb -- for $0.00 and under a second. It also beats the paid arm on the headline.
Half your evidence lives in addenda, peer-to-peer notes and a withdrawal somebody sent by fax — the paid arm It is the only thing here that reaches the correspondence channel -- the free floor takes 64.7 pct of those 34 grounds and the arm takes 91.2 pct, and on the return calls only the prose makes returnable the floor scores 0 of 6 against the arm's 6 of 6.
You want one answer and you are choosing between them — both, in that order -- code first, model on what is left They fail differently and the failures barely overlap. The floor is perfect on the tables and blind to the prose; the arm is 82.4 pct on the tables and 91.2 pct on the prose. Running the floor first costs nothing and leaves the model a smaller, harder job.
At a glanceHow the whole thing runs
86%across runs
180,268 msp50, end to end
$71.22per 1,000 prior-authorisation packets · Google Gemini 3 Flash
Run once, for real, on 2026-08-29. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check what a prior-authorization packet actually proves14 steps · 4 questions · run once, for real · 2026-08-29
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point src/app.py at your own packets by dropping .txt files into data/corpus/ in the same section layout -- underlined headings, a criteria table with the printed columns, and a Submitted Documents section. ⚠︎ DO NOT point this at a real packet.Corpus lens →
When is this the wrong choice?
Avoid: Paying for the part a spreadsheet already does. That is the case against the best-fitting scenario (“Your criteria table is clean and every document is filed beside it, correctly dated”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A REAL PACKET. src/packet.py is regular expressions written for this corpus's layout -- underlined headings, two-space indented tables, CR- / SD- / MP- / DK- identifiers. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether any of these rates resemble a real payer's queue. The corpus is invented and the defect mix is chosen; see data/SOURCES.md. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-29 — r001-priorauth-evidence. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 -m src.app. No key, no install, no index build and no network: requirements.txt is empty because the kit is standard library end to end. The corpus, the answer key and the recorded run all ship in the repo, so the whole product renders before anyone decides to spend.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
86.4%rows answered
180,268 msp50, end to end
274,015 msp95
5 minclone to first result
What the clock covers. one reading -- one prior-authorisation packet, end to end including provider-side reasoning tokens, on a shared connection. ⚠︎ NOT AN SLA AND NOT COMPARABLE ACROSS KITS: this credential is shared with sibling kits and these figures are a wall-clock observation, not a throughput measurement. 40 packets ran in 967.3 wall seconds with 8 concurrent workers; the free floors answer the same 40 packets in under a second with no network at all.
Current processWhat it replaces
The manual first-pass read of a submitted packet against the plan's criterion set, before clinical review.
Where it is not good enough
THE STRONGEST FREE FLOOR BEATS IT OVERALL AND THAT IS PUBLISHED HERE RATHER THAN BURIED. Pure Python (evidence-gate) scores 94.5 pct against the model's 86.4 pct, and settles 100 pct of the grounds the printed tables prove against the model's 82.4 pct, with zero false gaps and zero noise returns against the model's 6 and 11. The model earns its place on one thing only: the 91.2 pct of correspondence-settled grounds the floor reaches 64.7 pct of, and the 6 of 6 prose-only return calls the floor gets 0 of 6. Use code for the criteria table and a model for the correspondence.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt40jsonl1json1
40 prior-authorisation packets, 220 criteria — the plan's medical policy, the criteria table, the submitted documents, the clinical correspondence and the review notes
the rulebook is the packet's own Medical Policy And Evidence Terms — eight clauses, printed in every one of the 40 packets
src/recheck.py runs the structured checks off the printed columns, free: kind, lookback, minimum duration, MP-6.1
MP-2.1 is applied AS WRITTEN — a record stating a PERFORMED date is judged on that date, not on the day the report was issued
Recorded failurean earlier cut of the corpus planted duration_short on criteria whose table prints no minimum, 11 of 14, so the key asserted a defect the packet does not show — evals/check_labels.py named it before a single paid call
one entry per criterion: disposition, the ground of six, the clause this packet prints and the documents behind it
plus a SEPARATE return call, the evidence date copied off the document and the days short; 0 of 220 criteria were left off entirely
Recorded failure6 of the 42 satisfied criteria were called gaps anyway — most often an MP-5.1 not-applicable criterion read as missing evidence, which is a packet returned for something the policy never required
220 criteria, exact match against data/gold.jsonl, free, no judge grades anything
two headline numbers, never blended: 74 table grounds, 34 correspondence-only, 42 traps, 10 evidence gaps
0 invented citations of 220 — every identifier written appears in the packet in front of the arm
Recorded failurethe correspondence-blind ablation loses 53 points of correspondence-settled grounds, 91.2 to 38.2, and holds the table ones at 85.1 — that split is the evidence the prose was read rather than guessed around
return agreement 89.5 pct against the floor's 91.8
2026-08-29as of
A health plan's review team, one prior-authorisation packet at a time, before the clinical determination.
⚠︎ EVERYTHING HERE IS SYNTHETIC: every member, plan, medical policy, criterion, clinician, submitted document and piece of correspondence was invented by tools/build_corpus.py from seed 20260829, and the policy is a plausible composite that is nobody's policy. No real packet, medical record or payer criterion set was used, reproduced or approximated, and none could be — the member's record is protected health information and the enhanced-review list is a utilisation-management position no payer publishes. There is no PHI in this repository. The defect mix is CHOSEN, so no rate here estimates how often a real submission is incomplete. ⚑ THE NUMBER TO READ FIRST IS THE FREE FLOOR'S, AND ON THIS KIT IT BEATS THE MODEL: evidence-gate takes 94.5 pct of question one for $0.00, with no key and no network, in 0.1 s across all 40 packets, against the fast tier's 86.4 — the floor is 8.1 points AHEAD, and it is ahead on question two as well, 91.8 against 89.5. Pure Python settles every ground the printed tables prove, 74 of 74, where the arm takes 82.4; it raises 0 false gaps of 88 where the arm raises 6, and 0 noise returns of 112 where the arm returns 11. ⚑ SO WHAT THE MONEY BUYS IS NARROW AND IT IS REAL: the 34 grounds only the clinical correspondence settles, where the floor reaches 64.7 pct and the arm takes 91.2 — and the 6 return calls that are returnable ONLY because of something a note says, where the floor scores 0 of 6 and the arm 6 of 6. The honest conclusion is not 'use a model'. It is: use code for the criteria table and a model for the correspondence, in that order, because running the floor first costs nothing and leaves the model a smaller and harder job.
⚠︎ AND THE FLOOR'S 100 PCT ON THE PRINTED TABLES IS AN UPPER BOUND, NOT A FORECAST. It implements the same rules the generator used, so on a real packet — dates in prose, kinds implied rather than labelled, documents that do not declare which clause they were filed under — it would do worse. The arm's 82.4 pct on the same rows carries no such advantage, and that asymmetry is the single largest caveat on this page. ⚑ THE ABLATION IS THE EXPERIMENT, NOT A SECOND OPINION: removing the clinical correspondence costs the arm 53 points of correspondence-settled grounds, 91.2 to 38.2, takes its prose-only return calls from 6 of 6 to 0 of 6 — exactly the free floor's score — and leaves the table-settled grounds alone at 85.1, which is UP, not down. An arm pattern-matching on the criteria table would have lost both or neither. It also cost MORE per packet than the full arm despite a smaller prompt, because a model given less evidence reasons for longer; trimming context is not reliably a cost lever here.
⚠︎ THE CEILING WAS PROBED BEFORE THE SCORED RUN, NOT AFTER: at 32,000 tokens the calibration probe's heaviest packet drew 30,632, 95.7 pct of the cap, so the published ceiling was raised to 64,000 BEFORE the arm fired — and the scored run's largest reply then drew 40,963, 28 pct over the ceiling that was replaced. Choosing a ceiling after seeing a truncated run is choosing it after the game.
⚠︎ ONE RUN. Every arm here was fired once, and the provider re-rolls its reasoning budget per call, so the run-to-run spread is unmeasured.
⚠︎ NO INJECTION PROBE WAS RUN and none is claimed — which matters more here than on most kits, because the documents in a real packet are supplied by the party being checked. That is recorded under could_not_verify rather than left to look like an absence of risk.
⚠︎ IT PRODUCES AN EVIDENCE CHECK FOR A PERSON TO READ. It never approves, denies, authorises or declines a request — there is no such endpoint and no flag that adds one.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER + BASE_URL + MODEL in .env -- one line, then the same run again. Two adapters ship (an OpenAI-compatible shape and Anthropic's Messages API) and adding a third is one function and one entry in PROVIDERS. It must return token counts, because the Cost lens prices them.
what leaves this machine
src/select.py
NEVER_SENT -- the named sections withheld at the seam rather than trusted to a sentence in a prompt. The UI prints what went and what stayed, per packet, so a reader can check rather than believe.
the rules a free floor applies
src/recheck.py
The four switches on check() -- use_performed, refuse, honour_exception, apply_mp61. Each is a clause the policy states and a naive checker skips; turning them on one at a time is what separates the three floors.
the evaluation
evals/scoring.py
Graders and denominators. The two questions are scored on separate denominators and can be reweighted independently; a scorer that collapsed them would hide the finding the kit exists to show.
Components
Component
File
Role
packet parser
src/packet.py
One packet into structured data -- header, criteria table, submitted documents, correspondence. Strict about identifier shapes, and a row it cannot recognise is skipped rather than guessed at, because a mis-parsed criterion silently changes a denominator.
the seam that withholds
src/select.py
Drops the Plan Utilisation Position before anything is sent. Every figure in that section is a materiality position, and materiality is the second question this kit measures.
prompt assembly
src/prompt.py
System instruction plus the packet verbatim, minus the withheld section. Names the five dispositions, six grounds and two return calls; never says where to look.
the model call
src/adapters/__init__.py
Raw HTTP to any OpenAI-compatible provider or Anthropic. Retries a busy provider, refuses to retry a wrong request, and checks the shared daily call cap before spending.
the rules, in code
src/recheck.py
Lookback arithmetic, kind matching, duration checks and MP-6.1 -- the deterministic half, which the free floors and the UI both run.
the scorer
evals/scoring.py
Two headline numbers on separate denominators, per-channel and per-pattern breakdowns, and a fabrication check read from the packet rather than the key.
Where it breaks at scale
A real packet is a PDF bundle, not a printed table. This kit reads a structured packet and the extraction problem sits upstream of everything here, unsolved. Beyond that: one packet is read whole on every call, so a packet with 60 criteria and 200 documents would push the input past what a single call should carry, and the answer would need chunking by criterion -- which reintroduces the cross-document reasoning (an addendum withdrawing a record filed under a different criterion) that this kit's hardest cases turn on.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
PA-0014, replayed from the scored run. Six criteria and both arguments at once: CR-0014-05 is a filed, in-window imaging report of the right kind that a peer-to-peer addendum WITHDRAWS -- the free floor calls it met, the model calls it unevidenced under MP-3.1 and cites the addendum. CR-0014-02 is the second argument: both arms call it stale, the floor waives the return under MP-6.1's equivalent-document allowance and the model does not, because this service is on the enhanced-review list -- a fact recorded ONLY in the correspondence on this packet. 4 of 6 criteria are identical to the free floor; the two that are not are the two the kit exists for.successOpen full size →The same packet with NO API key configured. The criteria table, every document's kind and date, the window arithmetic, the free floor's answer on all six criteria and the seam table are all computed locally and render in full. The model column is the only thing missing, and the button says why rather than failing at the HTTP layer.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
PA-0011, the packet the scored run got MOST wrong -- four of its criteria. Replayed from the same run, so what is on screen is inside the published percentages rather than a second attempt that might come out differently. A kit that only ships the packets it got right is an advert.failureOpen full size →
No train/test split, because nothing is trained or tuned. 40 packets, each one request under one plan with the plan's medical policy, the criteria table, the submitted documents, the clinical correspondence and the review notes. Every criterion on every packet is labelled and every criterion is scored -- 220 cells. 88 are already met, 122 are genuine gaps and 10 cannot be settled from the packet at all. Of the gaps, 74 are proved by the packet's own printed tables and 34 are decided ONLY by the clinical correspondence. And 42 of the met criteria are TRAPS -- every one looks like a gap on the criteria table and every one is correct. The second question has its own split: 108 criteria send the packet back and 112 do not, and of those that do, 6 are returnable only because of something the correspondence says.
SetupWhat the setup figure measured
There is no index to build. The whole packet, minus the withheld Plan Utilisation Position, goes into the prompt verbatim -- no chunking, no retrieval, no pre-digest, at 2656 input tokens on average. The 0.1 s is what pure-code parsing plus the strongest free floor cost across all 40 packets.
LicenceLicence
MIT, with the corpus generated in-process -- there is no third-party data in this kit to licence.
Bring your ownBring your own prior-authorisation packets (one request each, 5-6 criteria)
Point src/app.py at your own packets by dropping .txt files into data/corpus/ in the same section layout -- underlined headings, a criteria table with the printed columns, and a Submitted Documents section. Everything downstream (the floors, the seam, the scorer) reads the parser's output, so a new corpus needs no code change. What it DOES need is a gold.jsonl of your own; without one the UI works and the eval has nothing to score against.
⚠︎ And what stops being true when you do: ⚠︎ DO NOT point this at a real packet. It carries PHI, this kit has no de-identification step, and src/select.py withholds a NAMED SECTION rather than redacting content -- a real record would be sent to your provider in full.
What breaks it
A REAL PACKET. src/packet.py is regular expressions written for this corpus's layout -- underlined headings, two-space indented tables, CR- / SD- / MP- / DK- identifiers. Pointed at a PDF bundle or a portal export it parses NOTHING, and parse() returns empty tables rather than guessing. It also carries PHI, which this kit has no step for -- see bring_your_own_boundary.
A CRITERION WHOSE NAMED DOCUMENT IS ABSENT. The parser reports it as missing rather than treating it as unevidenced -- 'not filed' and 'filed and insufficient' are different states and only one of them is a finding a practice can answer.
A DOCUMENT THAT DOES NOT DECLARE WHICH CLAUSE IT WAS FILED UNDER. MP-6.1 equivalents and MP-4.1 exception rationales are found here by reading the document's own text. A real packet offers no such convenience, and both classes would need judgement rather than a lookup.
A CRITERION CARRYING TWO DEFECTS AT ONCE. The scorer asks for exactly one ground per criterion, so a document that is both outside the lookback AND of the wrong kind has no single right answer. The generator places patterns on compatible criteria to keep that from happening; evals/check_labels.py is what would name it if it did.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
the task, the vocabularies and the output shape
5,937
1,253
the packet, verbatim, minus the withheld section
6,645
1,403
Total
2,656
This is the cost lesson as arithmetic: of the 2,656 tokens assembled, 1,403 are documents — 53% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
The exact system instruction and packet body sent for PA-0014 in run r001-priorauth-evidence. Assembled by src/prompt.py; the packet goes in verbatim minus the withheld section.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are checking one prior-authorisation packet for EVIDENCE, against the criterion set
the packet itself prints for the service requested, before any clinical review has happened.
You act for the PLAN'S REVIEW TEAM. This is not a determination and not a denial: it is the
evidence check a reviewer reads first, and what the requesting practice will answer criterion by
criterion. A gap you cannot attach to a clause is not something anyone has to answer at all, and a
packet returned for something the policy already allows is how this whole exercise gets switched
off.
Return one entry for EVERY criterion in the "Criteria" table, in the order the packet lists them,
using the packet's own criterion identifiers. Never drop a criterion because it looks unremarkable.
For each criterion give exactly one disposition:
CRITERION_MET the criterion is satisfied -- a document of the required kind, inside the
required window, states the fact the criterion requires. It is also met
where the policy in this packet says it is met another way, or says it
does not apply to this presentation.
CRITERION_UNEVIDENCED nothing filed with this packet establishes the fact -- there is no
document, or what is filed does not actually evidence it.
CRITERION_STALE a document establishes the fact and it falls outside the criterion's
lookback window.
CRITERION_WRONG_ARTEFACT the fact is asserted, but not by the kind of document the criterion
names.
EVIDENCE_INCOMPLETE something about this criterion cannot be settled from this packet -- the
document that would decide it is named and not filed. Do not state a
finding; say what is missing.
When the disposition is not CRITERION_MET and not EVIDENCE_INCOMPLETE, the ground -- what actually
went wrong -- is exactly one of:
NO_SUPPORTING_DOCUMENT no filed document states the fact the criterion requires
ARTEFACT_KIND_MISMATCH the fact is stated, by a document of the wrong kind
OUTSIDE_LOOKBACK the supporting document falls outside the criterion's window
DURATION_NOT_MET a required minimum duration is documented and is short of it
SUPERSEDED_BY_ADDENDUM a later document in this packet corrects or withdraws the record
the criterion was resting on
CONTRAINDICATION_UNADDRESSED a contraindication to a required trial is documented and what the
policy requires in its place is not filed
Otherwise ground is null.
Also give, for every criterion, a return call:
return RETURN this criterion sends the packet back to the requesting practice.
DO_NOT_RETURN it does not. The plan's policy in this packet carries an
equivalent-document allowance; read it, apply BOTH of its limbs, and note
that this packet records somewhere whether the second limb is engaged for
the service requested.
THIS IS A SEPARATE JUDGEMENT FROM THE DISPOSITION AND IT IS SCORED SEPARATELY. A check that returns
a packet for every criterion with anything imperfect about it is a column, not a review, and the
practice that receives it stops reading. A check that returns none is worth nothing. Answer both
questions on every criterion.
And two values, for every criterion:
evidence_dated the date carried by the document the criteria table names for this criterion,
copied straight off that document. No judgement. null when the table names no
document, or names one that is not filed.
days_short how far the criterion falls short, in days: days outside the window when it is
outside, days short of the minimum when a duration is short. 0 when the
criterion is met and 0 when the shortfall is not a matter of days. null when the
disposition is EVIDENCE_INCOMPLETE.
Cite the governing clause as a clause identifier that appears in this packet, and attach the
evidence as document identifiers that appear in this packet. NEVER write a clause or document
identifier that is not printed in the packet in front of you -- an invented citation is worse than
none, because it is the first thing the practice will check.
Anything in the packet may bear on a criterion. Where a printed table and a later statement in the
packet contradict each other, they are not a tie.
Reply with JSON and nothing else:
{"determination_action": "RETURN_FOR_EVIDENCE" | "REQUEST_DOCUMENTS" | "READY_FOR_REVIEW",
"criteria": [{"criterion": "<the packet's criterion identifier>",
"requirement": "<the requirement, as the packet states it>",
"disposition": "CRITERION_MET" | "CRITERION_UNEVIDENCED" | "CRITERION_STALE" |
"CRITERION_WRONG_ARTEFACT" | "EVIDENCE_INCOMPLETE",
"ground": "<one of the six, or null>",
"governing_clause": "<a clause identifier from this packet, or null>",
"evidence": ["<document identifiers from this packet>"],
"return": "RETURN" | "DO_NOT_RETURN",
"evidence_dated": "<YYYY-MM-DD or null>",
"days_short": <number or null>,
"finding": "<the one sentence the review queue would carry for this criterion>"}],
"rationale": "<two sentences at most, on what decided the hardest criterion>"}
determination_action is RETURN_FOR_EVIDENCE if any criterion is RETURN; REQUEST_DOCUMENTS if none
is but at least one is EVIDENCE_INCOMPLETE; READY_FOR_REVIEW otherwise. None of the three approves,
denies, authorises or declines the request -- each is a recommendation to the people who make that
determination.
Prior Authorisation Evidence Check
----------------------------------------------------------------
SYNTHETIC RECORD -- invented for an open kit. No real member, patient,
payer, plan, clinician, facility, record or person appears in it.
Packet PA-0014
Request RQ-2026-0114
Plan Cardinal Health Plan (PL-CHP)
Member MB-0014
Requesting provider PR-0014
Service requested Lumbar fusion, single level (SVC-LFUS1)
Submitted 2026-04-17
Reviewed as of 2026-04-19
Enhanced review not recorded on this header
Medical Policy And Evidence Terms
----------------------------------------------------------------
Evidence policy
MP-1.1 A criterion is met only where a document filed with this packet
states the fact the criterion requires.
MP-1.2 The filed document must be of the KIND the criterion names. A
statement in a progress note that a study was performed, or that
it supported the request, is not that study.
MP-2.1 Every criterion carries a lookback window, measured back from the
submission date. Evidence dated outside that window does not meet
the criterion, however clearly it states the fact. Where a record
states both a performed date and a report date, the PERFORMED date
is the one this clause runs on.
MP-2.2 Where a criterion carries a minimum duration, the filed evidence
must show that duration completed on or before the submission date.
MP-3.1 A later document that corrects or withdraws an earlier one governs
from its own date, however the criteria table still reads.
MP-4.1 Where a contraindication to a required trial is documented, the
criterion is met by a filed exception rationale instead. The
contraindication alone does not meet it.
MP-5.1 A criterion the policy makes inapplicable to the presentation is
not an evidence gap and is not returned.
MP-6.1 A criterion that is not met is NOT returned where an equivalent
document on file states the same fact, UNLESS the service is on
this plan's enhanced-review list for the current benefit year.
Document kind register
DK-EXC Exception rationale
DK-IMG Imaging report
DK-LAB Laboratory report
DK-MED Medication record
DK-PROG Progress note
DK-SPEC Specialist consultation
DK-THER Therapy note
Clause index
MP-1.1 Filed document required
MP-1.2 Document kind
MP-2.1 Lookback window
MP-2.2 Minimum duration
MP-3.1 Correction or withdrawal
MP-4.1 Contraindication exception
MP-5.1 Inapplicable criterion
MP-6.1 Equivalent document
Criteria
----------------------------------------------------------------
Criterion Requirement Kind Look Min Document
CR-0014-01 Tobacco cessation counselling recorded DK-PROG 6m -- SD-0014-01
CR-0014-02 Conservative therapy trial completed DK-THER 12m 42d SD-0014-02
CR-0014-03 Specialist consultation on file DK-SPEC 12m -- SD-0014-03
CR-0014-04 Neurological deficit documented on examination DK-PROG 6m -- SD-0014-04
CR-0014-05 Advanced imaging confirms the level treated DK-IMG 12m -- SD-0014-05
CR-0014-06 Neurological deficit documented on examination DK-PROG 6m -- SD-0014-06
Lookback windows are measured back from the submission date, 2026-04-17.
Submitted Documents
----------------------------------------------------------------
Document SD-0014-01
Kind Progress note (DK-PROG)
Dated 2025-04-09
Tobacco cessation counselling recorded.
Recorded at this encounter and filed with this packet.
Document SD-0014-02
Kind Therapy note (DK-THER)
Dated 2024-12-30
Conservative therapy trial completed.
Recorded at this encounter and filed with this packet.
Document SD-0014-02-E
Kind Therapy note (DK-THER)
Dated 2025-12-18
Conservative therapy trial completed.
This record states the same fact as SD-0014-02 and is filed under MP-6.1 as the equivalent document for criterion CR-0014-02.
Document SD-0014-03
Kind Specialist consultation (DK-SPEC)
Dated 2024-11-08
Specialist consultation on file.
Recorded at this encounter and filed with this packet.
Document SD-0014-04
Kind Progress note (DK-PROG)
Dated 2026-03-18
The member has a documented contraindication to the trial this criterion requires.
The trial required by this criterion was therefore not attempted.
Document SD-0014-04-X
Kind Exception rationale (DK-EXC)
Dated 2026-03-28
Exception rationale filed under MP-4.1 against criterion CR-0014-04.
The contraindication recorded in SD-0014-04 is accepted as the basis.
Requirement addressed by exception: Neurological deficit documented on examination.
Document SD-0014-05
Kind Imaging report (DK-IMG)
Dated 2026-03-15
Advanced imaging confirms the level treated.
Recorded at this encounter and filed with this packet.
Document SD-0014-06
Kind Progress note (DK-PROG)
Dated 2025-10-28
Neurological deficit documented on examination.
Recorded at this encounter and filed with this packet.
Clinical Correspondence
----------------------------------------------------------------
Addendum AD-0014-05, 2026-03-20, from the requesting practice: the finding recorded in SD-0014-05 is withdrawn. On review the study reported there was the member's prior study and does not support this request. SD-0014-05 should not be read as evidencing "advanced imaging confirms the level treated".
Utilisation management note, 2026-04-10: SVC-LFUS1 is on this plan's enhanced-review list for the current benefit year. MP-6.1's equivalent-document allowance does not apply to requests for this service.
Review Notes
----------------------------------------------------------------
Packet assembled by the requesting practice and checked against the criterion set printed above. This is an evidence check before clinical review; it is not a determination.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"determination_action":"READY_FOR_REVIEW","criteria":[{"criterion":"CR-0001-01","requirement":"Conservative therapy trial completed","disposition":"CRITERION_MET","ground":null,"governing_clause":"MP-2.2","evidence":["SD-0001-01"],"return":"DO_NOT_RETURN","evidence_dated":"2026-01-23","days_short":0,"finding":"SD-0001-01 documents a completed conservative therapy course of 49 days, within the 12-month window and meeting the 42-day minimum."},{"criterion":"CR-0001-02","requirement":"Neurological deficit documented on examination","disposition":"CRITERION_MET","ground":null,"governing_clause":"MP-2.1","evidence":["SD-0001-02"],"return":"DO_NOT_RETURN","evidence_dated":"2026-02-13","days_short":0,"finding":"SD-0001-02 is a progress note dated 2026-02-13 documenting the neurological deficit, within the 6-month window."},{"criterion":"CR-0001-03","requirement":"Tobacco cessation counselling recorded","disposition":"CRITERION_MET","ground":null,"governing_clause":"MP-6.1","evidence":["SD-0001-03","SD-0001-03-E"],"return":"DO_NOT_RETURN","evidence_dated":"2025-06-14","days_short":0,"finding":"SD-0001-03 is outside the 6-month window, but SD-0001-03-E, a same-kind progress note dated 2026-04-01, states the same fact and the header records no enhanced review, so MP-6.1 applies."},{"criterion":"CR-0001-04","requirement":"Specialist consultation on file","disposition":"CRITERION_MET","ground":null,"governing_clause":"MP-2.1","evidence":["SD-0001-04"],"return":"DO_NOT_RETURN","evidence_dated":"2026-01-30","days_short":0,"finding":"SD-0001-04 is a specialist consultation dated 2026-01-30, within the 12-month window."},{"criterion":"CR-0001-05","requirement":"Advanced imaging confirms the level treated","disposition":"CRITERION_MET","ground":null,"governing_clause":"MP-2.1","evidence":["SD-0001-05"],"return":"DO_NOT_RETURN","evidence_dated":"2026-01-06","days_short":0,"finding":"SD-0001-05 is an imaging report dated 2026-01-06 confirming the level treated, within the 12-month window."}],"rationale":"CR-0001-03 is the hardest criterion: SD-0001-03 is outside the 6-month window, but SD-0001-03-E is a same-kind progress note dated 2026-04-01 stating the same fact, and the header does not record enhanced review for this service, so MP-6.1 prevents a return. The remaining criteria are each satisfied by their named documents within the applicable windows and, for CR-0001-01, the required minimum duration."}
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check what a prior-authorization packet actually proves — 40 prior-authorisation packets drawn from 40 real prior-authorisation packets (one request each, 5-6 criteria). One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
No model grades anything. Every metric is an exact match against data/gold.jsonl, and the key itself is gated by evals/check_labels.py, which re-derives every label from what the packets print -- through the same parser the arms use -- rather than from the generator's intent. That gate convicted four real key defects before a single paid call.
40prior-authorisation packets
40source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED190 · 173 · 208 · 140 · 78 / 220evidence accuracy pct — criterion, defensible findingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED197 · 202 / 220return agreement pct — criterion, return callDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED31 · 13 · 22 / 34casefile ground caught pct — ground only the correspondence settlesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED61 · 63 / 74table ground caught pct — ground the printed tables proveDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
evals/check_labels.py re-derives every label FROM THE PACKETS -- through the same parser the arms use, not from the generator's intent -- over 220 criteria, and asserts eleven properties of the key: the packet and key agree criterion for criterion; every identifier the key cites is printed in its packet; the clause is the one the policy prints for that ground; the window and duration arithmetic reproduces days_short exactly; EVIDENCE_INCOMPLETE names a document that really is absent; and MP-6.1's allowance is never claimed on an enhanced-listed service. It convicted four real key defects before a single paid call and is green now.
Grading costWhat it costs
Every dollar on this page is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid. The token counts come from the provider's own usage block on each call.
Priced at
Per 1M in / out
One prior-authorisation packet
1,000 prior-authorisation packets
Share that is the prompt
Google Gemini 3 Flash the estate's shared projection card, so this kit's rows compare with every other kit's
$0.50 / $3.00
$0.071220
$71.22
2%
Same work, 1× the bill
The same prior-authorisation packets, the same tokens — only the rate card changed. And on that card about 2% of what you pay is the prompt this pipeline sends, not the answer it writes.
Turn the reasoning budget down. It is 96.4 pct of the output and the whole of the cost; the kit sends the provider default and records what it sent. That is a measurement this kit has not made, and it is the first thing to try before changing model.
Rates checked 2026-08-29. The provider that actually ran every call here is kept off this page per the series rule. The real spend is in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Scoring is pure Python over committed JSON: 220 criteria in under a second, no key, no network. The ARM costs money; grading it does not.
Where the answer had to be read from -- the printed tables, or the correspondence whether an arm is reading the criteria table or the prose. Splits every ground into the 84 the printed tables prove and the 34 only the clinical correspondence settles. This is the pair that separates the model from the free floor, and the ablation is what proves the split is real rather than a coincidence of difficulty.
$0.00
no
yes
full packet 91.2% · the ablation 38.2%
Citations that do not appear in the packet whether any clause or document identifier the arm wrote is absent from the packet in front of it. Read from the PACKET, not the key, so the same check would convict the answer key if the key ever cited something a packet does not print.
$0.00
no
yes
the fast tier 0.0% · the ablation 0.0%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
The ablation is the separability test and it passed cleanly. Removing the clinical correspondence costs the model 53.0 points of correspondence-settled grounds (91.2 pct -> 38.2 pct) and leaves the table-settled ones alone (82.4 pct -> 85.1 pct, which is up, not down). Prose-only return calls go 6 of 6 to 0 of 6 -- exactly the free floor's score. An arm pattern-matching on the criteria table would have lost both or neither.
Set limitationsWhat this set cannot show
The corpus is deliberately UNBALANCED and the scorer prints the majority baseline beside every headline for that reason. 88 of 220 criteria are already met, so the best single constant answer scores 40.0 pct on the five-way disposition; at packet level 28 of 40 requests carry a return, so the best constant action scores 70.0 pct. Both are published.
It means the headline is only readable against the majority baseline, which is why the scorer computes and prints it on every arm. It also means the per-class rates here are NOT the ones a real payer would see: the defect mix is chosen rather than observed, and 12 packets are return-free by construction so the three-way action has three real classes.
The specification
Twelve packets are dealt only non-returning patterns, so READY_FOR_REVIEW and REQUEST_DOCUMENTS are real classes rather than rounding error.
A packet holds at most one trap_enhanced, because the enhanced listing is a property of the SERVICE and two would be one fact printed twice.
No packet holds both halves of the MP-6.1 pair -- on an enhanced-listed service the allowance applies to no criterion on the packet, so the two together would make the key contradict the clause it prints.
A pattern is dealt only to a criterion that can express it: duration_short needs a minimum, wrong_artefact needs a required kind that is not already a progress note.
The specification was tried, and it convicted itself four times
evals/check_labels.py runs the spec above over 220 criteria, through the same parser the arms use rather than the generator's intent. Every constraint was added because its absence was caught: three packets carried both halves of the MP-6.1 pair, which makes the key contradict the clause it prints; eleven duration_short criteria sat on criteria whose table prints no minimum; four wrong_artefact criteria were not mismatches at all; and ten criteria filed two different documents under one identifier. All four were fixed before a single paid call, and the gate is green now. No separate PROBE set was authored -- rows is absent rather than 0, because an absent probe and a probe that found nothing are different facts.
$0.00 -- the generator and the label gate are pure Python. No attempt was made to match a real payer's class frequencies, because none is public. See data/SOURCES.md.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Your criteria table is clean and every document is filed beside it, correctly dated
the free floor (evals/baseline.py, evidence-gate)
It takes 100.0 pct of the grounds the printed tables prove, cites the clause, attaches the document and applies MP-6.1's first limb -- for $0.00 and under a second. It also beats the paid arm on the headline.
paying for the part a spreadsheet already does
Half your evidence lives in addenda, peer-to-peer notes and a withdrawal somebody sent by fax
the paid arm
It is the only thing here that reaches the correspondence channel -- the free floor takes 64.7 pct of those 34 grounds and the arm takes 91.2 pct, and on the return calls only the prose makes returnable the floor scores 0 of 6 against the arm's 6 of 6.
believing a table-only checker has read your packet
You want one answer and you are choosing between them
both, in that order -- code first, model on what is left
They fail differently and the failures barely overlap. The floor is perfect on the tables and blind to the prose; the arm is 82.4 pct on the tables and 91.2 pct on the prose. Running the floor first costs nothing and leaves the model a smaller, harder job.
treating this as a single-winner comparison, which is what the headline alone would suggest
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
SUPERSEDED_MISSED
Withdrawn record read as current
3
PA-0032 CR-0032-03: the correspondence withdraws the filed note and the arm still called the criterion met. This is the class the free floor loses entirely -- the model reaches 31 of 34 and the remaining 3 are all of this shape.
MP61_OVERREACH
Equivalent-document allowance applied where it does not reach
1
One criterion where the arm waived the return under MP-6.1 on a service that IS enhanced-listed. The clause has two limbs and this is the limb that bites; the free floor gets this class wrong 6 times out of 6 when the listing is recorded only in prose.
TRAP_FLAGGED
Satisfied criterion called a gap
6
6 of the 42 trap criteria were called gaps -- most often an MP-5.1 not-applicable criterion read as missing evidence, which is the over-return failure the second question exists to measure. The free floor's count here is 0.
GROUND_WRONG
Right finding, wrong ground
8
A criterion correctly called unevidenced but attributed to NO_SUPPORTING_DOCUMENT where the key says DURATION_NOT_MET -- the document is filed and states a duration, it is simply short. The finding is right and the sentence the practice receives is wrong…
What we could NOT verify
Whether any of these rates resemble a real payer's queue. The corpus is invented and the defect mix is chosen; see data/SOURCES.md.
Whether the free floor's 100 pct on table-settled grounds would survive a real packet. The floor implements the same rules the generator used, so that figure is an upper bound, not a forecast. The model's 82.4 pct on the same rows carries no such advantage.
Whether a second model would rank the same way. One paid model was measured, against three free floors and one ablation; this is a trade-off, not a survey of providers.
Whether the extraction step upstream -- a PDF bundle into a criteria table -- is tractable. It is not attempted here and nothing in these numbers speaks to it.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
2,655.5
23,297.4
180,268 ms
$0.071220
the ablation arm
2,632.7
24,956.4
195,839 ms
$0.076186
pure Python
0
0
0 ms
$0.000000
pure Python, no refusal
0
0
0 ms
$0.000000
what a review team does today
0
0
0 ms
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-29. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Answering vs gradingGrading this kit costs nothing. Answering it three times is the bill.
The scorer (evals/scoring.py), the label gate (evals/check_labels.py) and all three free floors are pure code and cost $0.00 to run against any result set -- there is no LLM judge anywhere in the grading path. The figure above is the token counts of the three PAID runs (the 40-packet scored arm, the 40-packet correspondence-blind ablation and the 3-packet ceiling probe) priced at the same projected card cost_per_query_usd uses, not a second larger spend. The ablation is more than a third of it, and it is what proves the prose was read.
Cost driversWhat actually moves the bill
Output tokens, overwhelmingly. 96.4 pct of them are provider-side reasoning left at the tier's default, and the reasoning budget is re-rolled per call.
The packet body, which is roughly half the input and grows with the number of criteria and documents.
The instruction, which is fixed at ~1,500 tokens and is the half that does NOT grow.
Your volumeWhat it costs at your volume
Linear. There is no index, no retrieval and no shared state -- each packet is one independent call -- so ten times the packets is ten times the cost and the same latency per packet. What does not scale is the socket: at 64,000 tokens a reply holds an unstreamed connection for up to ~491s, so concurrency is bounded by file descriptors and provider limits rather than by anything in this kit.
Where pricing changes shape
THE CEILING, AND THE PROBE ALREADY CONVICTED THE OLD ONE. At 32,000 tokens the calibration probe's heaviest packet drew 30,632 (95.7 pct). The published ceiling is 64000 and the scored run's largest reply drew 40,963 -- 28 pct over the ceiling that was replaced. You are billed for tokens DRAWN, not for the cap, so raising it is nearly free; a run discarded for truncation costs the whole run.
THE SOCKET TIMEOUT IS THE SAME SETTING WEARING A SECOND NAME. Completions are not streamed. The probe measured 130.4 output tokens per second, so a reply filling the published ceiling holds a silent socket for roughly 491 seconds; TIMEOUT_S is 1800. Raising one without the other turns a truncation defect into a transport defect the retry policy pays for twice.
PROVIDER-SIDE REASONING IS LEFT AT THE DEFAULT AND IS THE BILL. It is 96.4 pct of the output here. A card that prices it separately from completion moves the unit cost by roughly that share, and a per-query average would hide it entirely.
THE ABLATION COSTS MORE, NOT LESS. Removing the correspondence shrank the prompt and RAISED the bill per packet, because a model given less evidence reasons for longer. Trimming context is not reliably a cost lever on this task.
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the estate runs, so the figure compares with every sibling kit. The finding here is not about the model: pure Python beats it on the headline, and what it buys is the correspondence.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
2,656input tokens · this run
23,297output tokens
$0.071what it actually cost
per-packet average, this run
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$1.140
$1.140
$28.49
2026-09-12
gemini-3-flash
Google
$2.849
$2.849
$71.22
2026-09-18
gemini-3-8-flash
Google
$3.574
$3.574
$89.36
2026-09-18
llama-5
Meta
$4.093
$4.093
$102.33
2026-09-18
claude-haiku-4-5
Anthropic
$4.766
$4.766
$119.14
2026-09-12
grok-4-5
xAI
$5.804
$5.804
$145.10
2026-09-18
grok-4-6
xAI
$5.804
$5.804
$145.10
2026-09-18
claude-sonnet-5
Anthropic
$9.531
$9.531
$238.29
2026-09-12
gemini-3-1-pro
Google
$11.395
$11.395
$284.88
2026-09-18
gpt-5-6-terra
OpenAI
$11.395
$11.395
$284.88
2026-09-12
gpt-5-6-sol
OpenAI
$19.063
$19.063
$476.57
2026-09-12
claude-opus-4-8
Anthropic
$23.829
$23.829
$595.71
2026-09-12
claude-opus-5
Anthropic
$23.829
$23.829
$595.71
2026-09-12
claude-fable-5
Anthropic
$47.657
$47.657
$1191.43
2026-09-18
claude-fable-5-1
Anthropic
$47.657
$47.657
$1191.43
2026-09-18
gpt-6-astra
OpenAI
$47.657
$47.657
$1191.43
2026-09-17
Read this against the numbers above
Projection only -- no other model was actually called against this corpus.
96.4 pct of the output tokens are provider-side reasoning left at the tier's default. A model that reasons less, or more, moves every row below by more than its headline rate does -- which is why the cheapest card here is not automatically the cheapest run.
THE OUTPUT SIDE IS 8.8x THE INPUT SIDE on this workload (23,297 output tokens against 2,656 input). These rows are therefore almost entirely a bet on the OUTPUT rate; a card that is cheap on input and dear on output is worse here than its headline suggests.
THE FREE FLOOR COSTS NOTHING AND SCORES 94.5 PCT -- higher than the paid arm's 86.4. Every row below should be read against $0.00 and a BETTER headline, not against zero capability. What the money buys on this kit is narrow and named in the Eval lens: the 34 grounds only the clinical correspondence settles, and the prose-only return calls.
The ablation costs MORE per packet than the full arm ($0.076186 against $0.071220) despite a smaller prompt, because a model given less evidence reasons for longer. Removing context is not a cost saving.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Six modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Four of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/packet.pypacket parser
One packet into structured data -- header, criteria table, submitted documents, correspondence. Strict about identifier shapes, and a row it cannot recognise is skipped rather than guessed at, because a mis-parsed criterion silently changes a denominator.
src/packet.py
# Read one prior-authorisation packet into structured data. Pure code, no model.
MET = "CRITERION_MET"
UNEVIDENCED = "CRITERION_UNEVIDENCED"
STALE = "CRITERION_STALE"
WRONG_ARTEFACT = "CRITERION_WRONG_ARTEFACT"
INCOMPLETE = "EVIDENCE_INCOMPLETE"
DISPOSITIONS = (MET, UNEVIDENCED, STALE, WRONG_ARTEFACT, INCOMPLETE)
GROUNDS = ("NO_SUPPORTING_DOCUMENT", "ARTEFACT_KIND_MISMATCH", "OUTSIDE_LOOKBACK",
RETURN = "RETURN"
NO_RETURN = "DO_NOT_RETURN"
src/select.pythe seam that withholds — a swap seam
Drops the Plan Utilisation Position before anything is sent. Every figure in that section is a materiality position, and materiality is the second question this kit measures.
You change it to: NEVER_SENT -- the named sections withheld at the seam rather than trusted to a sentence in a prompt. The UI prints what went and what stayed, per packet, so a reader can check rather than believe.
src/select.py
# SEAM 2 -- what leaves this machine. Pure code, and deliberately a denylist of one.
NEVER_SENT = ("Plan Utilisation Position",)
def sent(sec_names):
def body(text, sections_fn):
src/prompt.pyprompt assembly
System instruction plus the packet verbatim, minus the withheld section. Names the five dispositions, six grounds and two return calls; never says where to look.
src/prompt.py
# The assembled prompt, in one place, in the order it is sent.
CORRESPONDENCE_HEADING = "Clinical Correspondence"
SYSTEM = """You are checking one prior-authorisation packet for EVIDENCE, against the criterion set
VERDICTS = ("CRITERION_MET", "CRITERION_UNEVIDENCED", "CRITERION_STALE",
VERDICT_MEANINGS = {
BLIND_SYSTEM_SUFFIX = """
def build(text, blind=False):
def _strip_correspondence(body):
def render(parts):
src/adapters/__init__.pythe model call — a swap seam
Raw HTTP to any OpenAI-compatible provider or Anthropic. Retries a busy provider, refuses to retry a wrong request, and checks the shared daily call cap before spending.
You change it to: PROVIDER + BASE_URL + MODEL in .env -- one line, then the same run again. Two adapters ship (an OpenAI-compatible shape and Anthropic's Messages API) and adding a third is one function and one entry in PROVIDERS. It must return token counts, because the Cost lens prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
TIMEOUT_S = 1800
TRANSPORT_RETRIES = 1
def _post(url, headers, payload, timeout=TIMEOUT_S):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
src/recheck.pythe rules, in code — a swap seam
Lookback arithmetic, kind matching, duration checks and MP-6.1 -- the deterministic half, which the free floors and the UI both run.
You change it to: The four switches on check() -- use_performed, refuse, honour_exception, apply_mp61. Each is a clause the policy states and a naive checker skips; turning them on one at a time is what separates the three floors.
src/recheck.py
# The rules, in pure code. No model, no key, no network.
DISPOSITION_OF = {
CLAUSE_OF = {
DUR = re.compile(r"Duration completed:\s*(\d+)\s*days", re.I)
def _d(iso):
def window_start(submitted_iso, months):
def duration_days(doc):
def is_contraindication(doc):
def check(parsed, crit, use_performed=False, refuse=False, honour_exception=False):
def evidence_dated(parsed, crit):
evals/scoring.pythe scorer — a swap seam
Two headline numbers on separate denominators, per-channel and per-pattern breakdowns, and a fabrication check read from the packet rather than the key.
You change it to: Graders and denominators. The two questions are scored on separate denominators and can be reweighted independently; a scorer that collapsed them would hide the finding the kit exists to show.
evals/scoring.py
# Grade one arm against the answer key. Deterministic -- no model judges anything here.
FINDING_DISPOSITIONS = (PK.UNEVIDENCED, PK.STALE, PK.WRONG_ARTEFACT)
def _pct(n, d):
def _rows(answer):
def _ret(r):
def score(records, golds, valid_ids):
Start hereThe shortest path into it
src/packet.pyOne packet into structured data -- header, criteria table, submitted documents, correspondence. Strict about identifier shapes, and a row it cannot recognise is skipped rather than guessed at, because a mis-parsed criterion silently changes a denominator.
src/select.pyDrops the Plan Utilisation Position before anything is sent. Every figure in that section is a materiality position, and materiality is the second question this kit measures. A swap seam.
src/prompt.pySystem instruction plus the packet verbatim, minus the withheld section. Names the five dispositions, six grounds and two return calls; never says where to look.
src/adapters/__init__.pyRaw HTTP to any OpenAI-compatible provider or Anthropic. Retries a busy provider, refuses to retry a wrong request, and checks the shared daily call cap before spending. A swap seam.
src/recheck.pyLookback arithmetic, kind matching, duration checks and MP-6.1 -- the deterministic half, which the free floors and the UI both run. A swap seam.
evals/scoring.pyTwo headline numbers on separate denominators, per-channel and per-pattern breakdowns, and a fabrication check read from the packet rather than the key. A swap seam.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2655 input and 23297 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
NOT MEASURED, AND THAT IS THE FINDING. No prompt-injection or adversarial run was made for this kit. Nothing below claims a defence; this section exists so the absence is a published fact with a denominator of zero rather than a blank space a reader would read as safety.
API_KEY is read from the shared <code><repo>/.env</code>, this kit's own <code>.env</code>, or the real environment, in that order. Both files are gitignored from the first commit; this repo has never held a credential. The key never leaves the machine and is never printed.
The experimentWe attacked it — two gates
An indirect prompt injection has to clear two gates, and they behave differently. Both were run for real on 2026-08-29.
Boundary checked
What could go wrong
What has actually been measured
Whether an instruction written into a SUBMITTED DOCUMENT can talk the arm out of a gap it would otherwise raise
Counting where the corpus happens to contain such a sentence gives a denominator of nothing -- the packets are generated by tools/build_corpus.py and carry no adversarial content by construction. A run over them would report perfect resistance while never having attacked anything.
NOT BUILT. The probe this needs would replace a submitted document's prose with an instruction and fire only at packets the key says carry a gap, so the denominator is packets that had something to suppress. It does not exist, so there is no figure here and none is implied.
Whether the PACKET-LEVEL return call can be flipped while every criterion row is left alone
Scoring only the criterion cells. An attack that leaves all 220 rows intact and downgrades a packet from return to clear has done the same damage to the review queue, and a cell-only measure would score it as a perfect defence.
NOT BUILT. The two denominators are already kept apart by the scorer -- 220 criteria and the packet-level return call are separate numbers throughout -- so the measurement would slot in. Nobody has run it.
Both gates are written down because the shape of the measurement is the part that is easy to get wrong, and writing it before running it is what stops a future run reporting a padded denominator as resistance.
The result0 attacks fired, 0 held. The honest pair of zeros, published rather than omitted -- and 34 of 220 criteria carry the exposure.
0adversarial trials fired
0phrasings written in the repo
220criteria whose robustness is UNMEASURED
34of those, decided only by the correspondence -- the attack surface
NO TRIAL WAS RUN, AND THESE ARE THE FIGURES FOR THAT. There is no evals/injection.py in this kit, no phrasing was written and nothing was fired, so there is no suppression rate, no paired denominator and no scope to report. The arithmetic that would make one: an instruction planted as the closing paragraph of a submitted document on packets the key says carry a gap, every criterion paired against this same model's own un-injected answer in r001-priorauth-evidence, and both denominators -- the 220 criteria and the packet-level return call -- kept apart, which evals/scoring.py already does. All 220 criteria above are unmeasured against that, and the 34 decided only by the clinical correspondence are the ones a forced instruction would most plausibly move: they are the kit's measured premium over free code (91.2 pct against the floor's 64.7) and its attack surface, and those are the same 34 rows.
Read this twice
⚠︎ THE SUBMITTED DOCUMENTS REACH THE MODEL verbatim, and in this workflow they are supplied by the party being checked. That is the difference between this kit and one reading its own company's files: the attack surface is the input, and the input is adversarial by nature of the process, not by accident. It has not been tested here. Treat the 86.4 pct as a figure measured on a cooperative corpus, and treat an injection run as the first thing to build before this goes anywhere near a real queue.
HonestyWhat this does not prove
Everything on this page. No attack was executed, so every row is a description of a measurement that has not been made.
Whether the deterministic floor is a mitigation. b-evidence-gate reads only the printed tables and never reads free prose, so an instruction hidden in prose cannot reach it -- but that is an argument from its design, not a measured result.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
Never approve a request, never deny one, never write to the authorisation system, never instruct a clinician, and never tell a member what their plan will cover. This kit says which criteria a filed document actually evidences, and hands that to the reviewer who decides.
Stated on the UI, in the README and in src/app.py's docstring, and enforced by there being no such endpoint: the server exposes GET routes plus one POST that returns an evidence sheet, and nothing else. There is no PUT, PATCH or DELETE.
EvidenceDoes it hold?
What
Measured
No citation is invented
0 fabricated identifiers across 220 criteria on r001-priorauth-evidence. The fabrication grader checks every clause identifier and document reference the arm returns against what the PACKET prints, not against the answer key -- so it can convict the key too.
The withheld section stays withheld
The Plan Utilisation Position is removed by src/select.py before assembly. Every figure in it is a materiality position, and no arm ever saw one.
A ground the packet does not settle is not asserted
The ablation is the proof this is read rather than guessed: strip the clinical correspondence and casefile grounds collapse from 91.2 pct to 38.2 pct of 34, while table-settled grounds HOLD at 85.1 pct of 74. An arm that was guessing would have scored the same in both arms.
Every criterion comes back
220 of 220 criteria returned across 40 packets; 6 of 7 packet sections sent on 40 of 40, and the UI prints the split per packet.
The evidence date is copied, not recalled
97.4 pct of 190 dates copied correctly off the document. The free floor scores 100.0 pct on the same rows, which is the honest way round: reading a printed date is code's job, not a model's.
The limitWhat a guardrail is not
This is ABSENCE OF A WRITE PATH plus a prompt rule, not a runtime enforcement layer. There is no policy engine and nothing would stop a fork from adding one.
⚠︎ IT IS NOT AN INJECTION DEFENCE, AND NOTHING HERE MEASURES ONE. No adversarial run was made for this kit. A real prior-authorisation packet is assembled by the party being checked, which is precisely where an instruction-shaped sentence would arrive; that is unmeasured and is recorded as unmeasured in the Threat model.
The free floors are not a guardrail either -- they are a comparison. Nothing in the kit refuses to call the model on a criterion the deterministic rules already settled, which is exactly the cost lever the Cost lens names.
A defensible finding is not a correct one. 86.4 pct is agreement with a key derived from an INVENTED corpus; it is not evidence about a real payer's queue.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 34 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
28 measured by the latest run6 need the model half
Metric
Owner
Role
Why this one
priorauth-evidence-accuracy
The whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
alarm
question one against the strongest FREE floor's 94.5 pct, which is the bar and which the model currently does NOT clear; question two against the same floor's 91.8 pct -- separate denominators, watched separately; the false-gap count, which is 0 for the floor and 6 for the model — alarm on evidence_accuracy_pct falling further below the free floor's 94.5 pct -- the model already loses on the headline, so the thing worth alarming on is the ONE column it wins: casefile_ground_caught_pct falling toward the floor's 64.7 pct, at which point paying for a model buys nothing on this corpus.
priorauth-evidence-channel
Where the answer had to be read from -- the printed tables, or the correspondence
alarm
the gap between table_ground_caught_pct and casefile_ground_caught_pct -- that gap IS the argument for paying for a model; the ablation's casefile figure, which is what proves the prose was read rather than guessed around — alarm on the ablation's casefile_ground_caught_pct rising toward the full arm's. If removing the correspondence stops costing anything, the arm was never reading it and the kit's central claim is gone.
priorauth-evidence-fabrication
Citations that do not appear in the packet
alarm
any non-zero invented_citation_pct at all -- this is the one column with a right answer of exactly zero — alarm on a single invented identifier. An invented citation is worse than none, because it is the first thing the requesting practice will check, and it is the failure that would end trust in the whole output.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
40
different corpus — nothing is comparable
corpus.bytes
235,380
prior-authorisation packets edited — the count held, the bytes did not
split.count
220
the criterion count moved — a different set was scored
split.size_p50
5,914
the median size of one criterion moved
split.size_p95
6,700
the 95th-percentile size of one criterion moved
dataset.rows
40
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.1
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (action_cells 40, blind False, casefile_ground_cells 34, cells 220, clause_cited_cells 122, dataset_version priorauth-evidence-v1-40packets, dated_cells 190, days_short_cells 210, document_cells 190, finding_criteria 122, ground_named_cells 122, immaterial_cells 14, incomplete_cells 10, met_criteria 88, noreturn_cells 112, packets 40, packets_answered 40, return_agree_cells 220, return_cells 108, return_prose_only_cells 6, table_ground_cells 74, trap_cells 42, whole_finding_cells 220) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Ground the printed tables prove
82.4 pct
74 grounds settled by the packet's own tables
r001-priorauth-evidence exact match against data/gold.jsonl; the free floor b-evidence-gate scores 100.0 pct on the same rows.
Ground only the correspondence settles
91.2 pct
34 grounds no table can settle
r001-priorauth-evidence against b-evidence-gate's 64.7 pct and the correspondence-blind ablation's 38.2 pct on the same 34 rows.
Satisfied criterion wrongly called a gap
14.3 pct (lower is better)
42 satisfied criteria
r001-priorauth-evidence; free floor 0.0 pct on the same 42.
Noise returned to the desk
9.8 pct (lower is better)
112 return calls
r001-priorauth-evidence; free floor 0.0 pct on the same 112.
Evidence date copied off the document
97.4 pct
190 dated documents
r001-priorauth-evidence; free floor 100.0 pct on the same 190.
latency
not yet known
every model call in the run
A band is the spread between repeats, and this kit has a single run of record, r001-priorauth-evidence, so any ceiling stated here would be invented rather than measured. The p50 and p95 on the board are that run's own, read from its captured record in build/measured/runs/.
Input tokens, whole run
106,220 on r001-priorauth-evidence
the whole run
One run of record, so no repeat spread exists yet; the figure is r001-priorauth-evidence's own, read from its captured record in build/measured/runs/. It is the bill, stated so a rewrite that quietly changes it is visible, and deliberately not a target.
Output tokens, whole run
931,897 on r001-priorauth-evidence
the whole run
One run of record, so no repeat spread exists yet; the figure is r001-priorauth-evidence's own, read from its captured record in build/measured/runs/. It is the bill, stated so a rewrite that quietly changes it is visible, and deliberately not a target.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
comply · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b-date-sweep 2026-08-29
b-evidence-gate 2026-08-29
b-index-tieout 2026-08-29
casefile ground caught, %
35.3
64.7
0.0
clause cited, %
82.0
90.2
0.0
criteria omitted, %
0.0
0.0
0.0
criterion disposition accuracy, %
76.4
94.5
49.1
days short accuracy, %
94.3
100.0
70.5
determination action accuracy, %
70.0
100.0
70.0
document attached, %
81.1
100.0
81.1
evidence accuracy, %
63.6
94.5
35.5
evidence date, %
100.0
100.0
100.0
false finding rate, %
22.7
0.0
0.0
false findings
20
0
0
false gap, %
47.6
0.0
0.0
finding caught, %
90.2
90.2
16.4
ground named, %
82.0
90.2
16.4
immaterial returned, %
100.0
0.0
0.0
incomplete recognised, %
0.0
100.0
0.0
input tokens, whole run
0
0
0
invented citation, %
0.0
0.0
0.0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
majority criterion disposition, %
40.0
40.0
40.0
majority determination action, %
70.0
70.0
70.0
noise return, %
39.3
0.0
8.9
output tokens, whole run
0
0
0
return agreement, %
74.5
91.8
55.5
return caught, %
88.9
83.3
18.5
return prose only, %
100.0
0.0
0.0
table ground caught, %
100.0
100.0
27.0
not a time series No two of these 3 runs measured the same system — they differ on criteria_citing_any_id, floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
comply · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c000-priorauth-evidence-calibration 2026-08-29
r001-priorauth-evidence 2026-08-29
s001-priorauth-evidence-corr-blind 2026-08-29
casefile ground caught, %
87.5
91.2
38.2
clause cited, %
92.9
83.6
73.0
criteria omitted, %
0.0
0.0
0.0
criterion disposition accuracy, %
94.4
89.5
81.4
days short accuracy, %
83.3
92.4
90.0
determination action accuracy, %
100.0
87.5
82.5
document attached, %
100.0
100.0
96.3
evidence accuracy, %
88.9
86.4
78.6
evidence date, %
100.0
97.4
94.2
false finding rate, %
25.0
6.8
8.0
false findings
1
6
7
false gap, %
50.0
14.3
16.7
finding caught, %
100.0
89.3
75.4
ground named, %
92.9
82.8
68.0
immaterial returned, %
—
7.1
14.3
incomplete recognised, %
—
90.0
100.0
input tokens, whole run
8558
106220
105308
invented citation, %
0.0
0.0
0.0
model latency p50 ms
153644.00
180268.00
195839.00
model latency p95 ms
242740.00
274015.00
290976.00
majority criterion disposition, %
33.3
40.0
40.0
majority determination action, %
100.0
70.0
70.0
noise return, %
25.0
9.8
9.8
output tokens, whole run
67139
931897
998256
return agreement, %
94.4
89.5
79.5
return caught, %
100.0
88.9
68.5
return prose only, %
100.0
100.0
0.0
table ground caught, %
100.0
82.4
85.1
not a time series No two of these 3 runs measured the same system — they differ on action_cells, blind, casefile_ground_cells, cells, clause_cited_cells, criteria_citing_any_id, dated_cells, days_short_cells, document_cells, finding_criteria, ground_named_cells, immaterial_cells, incomplete_cells, max_tokens, met_criteria, noreturn_cells, output_tokens_max, packets, packets_answered, reasoning_tokens_total, return_agree_cells, return_cells, return_prose_only_cells, table_ground_cells, trap_cells, whole_finding_cells — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-priorauth-evidence-stub 2026-08-29
casefile ground caught, %
0.0
clause cited, %
0.0
criteria omitted, %
0.0
criterion disposition accuracy, %
49.1
days short accuracy, %
70.5
determination action accuracy, %
70.0
document attached, %
81.1
evidence accuracy, %
35.5
evidence date, %
100.0
false finding rate, %
0.0
false findings
0
false gap, %
0.0
finding caught, %
16.4
ground named, %
16.4
immaterial returned, %
0.0
incomplete recognised, %
0.0
input tokens, whole run
58281
invented citation, %
0.0
model latency p50 ms
0.00
model latency p95 ms
0.00
majority criterion disposition, %
40.0
majority determination action, %
70.0
noise return, %
8.9
output tokens, whole run
18739
return agreement, %
55.5
return caught, %
18.5
return prose only, %
0.0
table ground caught, %
27.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 28 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
the ground on one criterion
the packet's return call. A criterion wrongly called a gap adds a return; a gap wrongly called satisfied removes the only reason the packet was going back. One cell moves what lands on the reviewer's desk, not one row of a sheet.
measured
evidence_accuracy_pct 86.4 against return_agreement_pct 89.5 on the same 220 criteria -- the two denominators move together and are scored apart precisely so this cannot hide.
the clinical correspondence
every ground no table settles -- and nothing else. Removing it costs 52.9 points on those 34 rows and leaves table-settled grounds untouched.
measured
s001-priorauth-evidence-corr-blind: casefile 91.2 -> 38.2 of 34, table 82.4 -> 85.1 of 74. The ablation is the kit's central claim.
the reasoning budget
the bill, not the answer. 96.4 pct of output tokens are provider-side reasoning, and the arm given LESS evidence reasoned for LONGER.
measured
$0.076186 per packet on the ablation against $0.071220 on the full arm, with a smaller prompt.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
Ground the printed tables prove
⚠︎ THE BAND THE MODEL PAYS WITH. Free code is perfect here and the paid arm is not. Any drop below the floor on these rows means money is buying a worse answer than a table lookup.
Ground only the correspondence settles
⚑ THE BAND THAT IS THE ARGUMENT. This is the only column where the paid arm clearly wins, and the ablation is what proves the win is reading rather than luck. If this falls toward 64.7 the kit's case is gone.
Satisfied criterion wrongly called a gap
Rising at all. A false gap is a request bounced back to a clinician who already filed the document.
Noise returned to the desk
Rising at all. Every point is a reviewer opening a packet that did not need it.
Evidence date copied off the document
Falling. A wrong date is what turns a current study into a stale one and flips the disposition with everything else correct.
latency
nothing yet.
NextThe three you would add first
⚑ GATE THE CALL ON THE ONLY THING IT BUYSSend a packet to the model only where a criterion turns on the clinical correspondence. On table-settled grounds free code scores 100.0 pct against the arm's 82.4, and on the headline the floor beats the arm outright, 94.5 to 86.4. The paid arm earns its keep on 34 rows out of 220. Routing on that alone is the accuracy lever and the cost lever at once.
STOP THE ARM CALLING A SATISFIED CRITERION A GAP6 of 42 satisfied criteria were returned as gaps -- 14.3 pct, against the floor's 0.0. A false gap sends a request back to a clinician who already filed the document, which is the most expensive error in this workflow and the one a reviewer is least likely to catch.
CAP THE NOISE RETURNED TO THE DESK11 of 112 returns were noise -- 9.8 pct, against the floor's 0.0. Every one is a reviewer opening a packet that did not need opening.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
The answer-key gate (evals/check_labels.py) is free and runs before any spend. All three free floors are free and re-run in seconds. Only the paid arm and the ablation cost money, and the ablation is more than a third of the bill -- it is also the thing that proves the claim, so it is not the place to economise.
What this cannot tell you
Whether any of these bands resemble a real payer's queue. The corpus is invented and the defect mix is chosen; see data/SOURCES.md.
Whether the guardrail holds against an instruction written into a submitted document. No adversarial run was made -- see the Threat model, which records this rather than leaving it to look like an absence of risk.
Whether a second model would sit in the same bands. One paid model was measured.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
NO FRAMEWORK. A folder of readable Python, standard library end to end -- no orchestration layer, no vendor SDK, no retrieval stack, no pip install. requirements.txt names nothing and says why.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the model call
src/adapters/__init__.py
an LLM wrapper / Runnable
Raw HTTP over urllib to any OpenAI-compatible provider or Anthropic. A wrapper would buy retries, a provider registry and streaming; the first two are already here in a few dozen lines, and this kit never streams -- it makes one call per packet and reads the whole reply.
the packet parse
src/packet.py
a document / schema library
Hand-written parsing into header, criteria table, submitted documents and correspondence. It is deliberately STRICT: a row it cannot recognise is skipped rather than guessed at, because a mis-parsed criterion silently changes a denominator and every rate on this page has one.
the withholding seam
src/select.py
a redaction or policy layer
Drops the Plan Utilisation Position before anything is sent. This is the one seam a framework would most plausibly own, and it is nine lines: the section is named, it is removed, and the eval measures what the arm does without it.
the rules
src/recheck.py
a rules engine
Lookback arithmetic, document-kind matching, duration checks and MP-6.1. The clauses that matter are printed in the PACKET, not configured here -- which is why there is no policy file to go stale against the request it is applied to.
the run loop
evals/run.py
an orchestration / DAG framework
ThreadPoolExecutor over independent packets at 8 workers, each answer appended to a partial file before anything else happens to it. There is no graph to build: no packet depends on any other.
the scorer
evals/scoring.py
an eval framework / LLM judge
Exact match in pure Python against data/gold.jsonl, on two separate denominators. No model grades anything, which is why one eval pass costs $0.00 and takes under a second.
the local board
src/app.py
a web framework and a front-end build
http.server, hand-written HTML and one JS file. It renders with no key and replays a committed run for free, which is the whole requirement.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Linear, and concurrent only at the packet level. There is no agent, no tool loop, no retrieval and no state carried between packets; every packet is independent, which is exactly why 8 workers is the entire concurrency story.
The other sideWhat a framework costs you
Adding a provider means editing the PROVIDERS dict by hand. In exchange pip install pulls nothing and a forker runs this on whichever key they already hold.
There is no structured-output library, so the reply is parsed by hand. That is a real exposure and it is NOT retired by this run: the parse held across all three paid arms, but a stricter schema library would fail loudly where this fails quietly.
No retrieval framework, because there is no retrieval. The rulebook is the packet's own policy section read whole into the prompt -- 2,656 input tokens, about half of them the instruction. A kit whose rulebook outgrew a prompt would need the layer this one does not.
What we could NOT verify
Whether a structured-output library would have changed the parse rate. The reply shape held on every paid arm, so there is no failure here to attribute to its absence.
Whether an orchestration framework would help at a packet count this kit has never run. 40 packets at 8 workers took 967 wall seconds; nothing about that shape was stressed.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-priorauth-evidence on the fast tier, 2026-08-29. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
180,268 ms
not yet known
nothing yet.
Model, p95
274,015 ms
not yet known
nothing yet.
Input tokens
106,220
106,220 on r001-priorauth-evidence
—
Output tokens
931,897
931,897 on r001-priorauth-evidence
—
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-priorauth-evidence-calibration153,644 ms
r001-priorauth-evidence180,268 ms
s001-priorauth-evidence-corr-blind195,839 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b-date-sweep, b-evidence-gate, b-index-tieout recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the rulebook is data read whole into the prompt.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
no dated run record — see the Eval lens for what was measured
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
prior-authorisation packets
data/corpus/PA-<n>.txt — 40 files, 235380 bytes, generated from seed 20260829
6 of the 7 sections go to the provider; Plan Utilisation Position never does
the answer key
data/gold.jsonl — 220 criteria, one row per packet
never. It is read only by the scorer, after the arm has answered.
the recorded runs
results/eval-*.json — the scored arm, the ablation, the ceiling probe and three free floors
never. The UI replays them locally so a reader with no key sees the measured arm.
the shared call ledger
<repo>/.calls-ledger.jsonl — one line per live call, written BEFORE the call
never. It is gitignored, and it is the only place the real spend is recorded.
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 78
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
Runs on
Python 3 standard library only. requirements.txt names nothing. Node is needed for the screenshots and for nothing else.
The key
API_KEY is read from the shared <code><repo>/.env</code>, this kit's own <code>.env</code>, or the real environment, in that order. Both files are gitignored from the first commit; this repo has never held a credential. The key never leaves the machine and is never printed.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
rulebook
the plan's medical policy, read out of the packet itself and sent whole -- the eight clauses with their identifiers, the document-kind register, the clause index and the criterion set. MP-6.1's equivalent-document allowance travels with it, because a check that could not read it would be checking against a policy it has not been shown.
2656 input tokens per packet on the scored run, of which about 1253 are the instruction (estimated at four characters per token; the total is the provider's own count). (r001-priorauth-evidence)
A packet with 60 criteria and 200 documents would push the input past what one call should carry. The answer would need chunking by criterion, which reintroduces the cross-document reasoning -- an addendum withdrawing a record filed under a DIFFERENT criterion -- that this kit's hardest cases turn on.
Editing the policy text after a run. Every published rate is against the clauses as printed in the corpus, and rewriting one after reading the misses is choosing the scoreboard after the game.
model
one completion call per packet, carrying six of the packet's seven sections. The Plan Utilisation Position -- the plan's cost for the service, the reviewer's discretion budget and the return-rate target -- is removed by src/select.py BEFORE anything leaves the machine, because every figure in it is a materiality position and materiality is the second question this kit scores. No chunking and no pre-digest: a summary is a place for the operative document to be lost before the model ever sees it.
40 packets, one call each, 967.3 wall seconds at 8 concurrent workers; p50 180268 ms, p95 274015 ms. 6 of 7 sections sent on 40 of 40 packets, and the UI prints the split per packet. (r001-priorauth-evidence)
At the published 64000-token ceiling a reply holds a silent, unstreamed socket for up to ~491 s, so concurrency is bounded by file descriptors and provider limits rather than by anything in this kit. The seam withholds a NAMED SECTION and is not a redaction system: a packet mentioning the return-rate target inside the correspondence would send it.
Sending the withheld section. The arm would be reading the answer to its own hardest question off the page, and the resulting score would be published as a judgement.
labels
nothing. data/gold.jsonl is read only by the scorer, after the arm has answered -- 220 criteria across 40 packets, every one labelled with its disposition, ground, clause, evidence, return call and both numbers.
220 criteria scored, 0 left off. Scoring is exact match in pure Python: no model grades anything and the whole set takes under a second. (evals/check_labels.py)
The key stops scoring where the corpus stops: it says nothing about a real packet, and its class frequencies are chosen rather than observed. It also assumes exactly ONE ground per criterion, so a document that is both outside the lookback and of the wrong kind would have no right answer.
Regenerating the corpus without re-running check_labels.py. The corpus and the key come out of one generator, so a bug there writes the packet AND the label that grades it, and the two agree with each other while both are wrong -- which is exactly how four real defects got in and were caught.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a criterion called unevidenced with no ground and no document attached
the arm has produced a column rather than a position. It is not something the requesting practice can answer, and it is scored as a miss on the whole finding even though the disposition may be right.
read ground_named_pct and document_attached_pct, not the headline (evals/scoring.py)
the free floor and the model agreeing on every criterion of a packet
the model bought nothing on that packet. On this corpus that is the COMMON case -- the floor is ahead overall -- and the UI says so per row rather than taking the credit.
open a packet whose enhanced listing is recorded only in the correspondence (the UI's agreement column)
invented_citation_pct above zero
the arm has cited a clause or document that is not printed in the packet. This is the one column whose right answer is exactly zero, and it is the first thing a practice would check.
stop. A single invented identifier ends trust in the whole output. (evals/scoring.py, read from the packet rather than the key)
output_tokens_max approaching max_tokens
the ceiling is about to bite. A truncated reply is recorded as a failure and stays in the denominator; it is not spliced out.
re-probe under a c run id before spending on a scored arm (results/eval-*.json, output_tokens_max)
A second paid tier through the model seam -- one model was measured, against three free floors and one ablation. Repeatability: every arm ran ONCE and 96.4 pct of the output is provider-side reasoning the provider re-rolls per call, so the run-to-run spread is unknown. A prompt-injection probe: not run, and it matters more here than on most kits, because the documents in a real packet are supplied by the party being checked. And the extraction step upstream -- a PDF bundle into a criteria table -- is not attempted at all.
The corpus licence, from the Data lens: MIT, with the corpus generated in-process -- there is no third-party data in this kit to licence. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineThe whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
whether each of the 220 criteria got a finding somebody could put in front of the requesting practice: the disposition the key carries, and on a gap the right one of six grounds, the clause the plan's own policy prints for it, every document the ground needs, and nothing cited that the packet does not print. And, SEPARATELY, whether the criterion sends the packet back at all under MP-6.1.
$0.00per 1,000 prior-authorisation packets
nodata leaves your network
yessame answer every time
MethodHow the test was run
python3 -m evals.run --run-id <id> [--floor <mode>|--blind]; evals/scoring.py then compares strings, dates and integers against data/gold.jsonl. No model is in the grading path.
Every grader on these pages scored the same 40 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
CR-0014-05
CR-0014-05
doc id
PA-0014
case
superseded
printed tables
every figure on this criterion's row checks out: SD-0014-05 is filed, dated 2026-03-15, inside the 12-month lookback from 2026-04-17, and is an Imaging report -- exactly the kind the criterion names
clinical correspondence
Addendum AD-0014-05: the finding recorded in SD-0014-05 is withdrawn. On review the study reported there was the member's prior study and does not support this request.
free floor
CRITERION_MET -- CR-0014-05 is met by the document filed against it.
model
CRITERION_UNEVIDENCED / SUPERSEDED_BY_ADDENDUM, clause MP-3.1, evidence SD-0014-05 and AD-0014-05, RETURN -- AD-0014-05 withdraws SD-0014-05, so no imaging report currently confirms the level treated.
The whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
pass
The key scores this criterion CRITERION_UNEVIDENCED, ground SUPERSEDED_BY_ADDENDUM, clause MP-3.1, return RETURN. The paid arm answered all four, so the criterion counts on both denominators — the finding and, separately, the return call. The free floor answered CRITERION_MET and loses the row: it reads the criteria table, where SD-0014-05 is filed, correctly dated and of the required kind, and nothing in that table records that a later addendum withdrew it. This is one of the 34 grounds only the correspondence settles.
Where the answer had to be read from -- the printed tables, or the correspondence
pass
Scored as a CASEFILE ground, not a table ground, because the addendum is the only thing that decides it — every figure the printed tables carry on this row checks out. It sits in the 34-row denominator where the arm scores 91.2 pct against the floor's 64.7, and the correspondence-blind ablation answers CRITERION_MET here, which is what the 38.2 pct is made of.
Citations that do not appear in the packet
pass
Both identifiers the arm cited — SD-0014-05 and AD-0014-05 — are printed in packet PA-0014, and MP-3.1 is a clause the packet's own policy section carries. Nothing was invented. This grader's right answer is exactly zero and the scored run returned 0 fabricated identifiers across all 220 criteria.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 86.4%
pure Python
scored 94.5%
In operationWhat to monitor
Reference standard: data/gold.jsonl, generated with the SYNTHETIC corpus and re-derived FROM IT by evals/check_labels.py. Every member, plan, policy, document and piece of correspondence is invented -- see data/SOURCES.md.
No true/false rates for this grader. It records 5 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
question one against the strongest FREE floor's 94.5 pct, which is the bar and which the model currently does NOT clear
question two against the same floor's 91.8 pct -- separate denominators, watched separately
the false-gap count, which is 0 for the floor and 6 for the model
Alarm on
evidence_accuracy_pct falling further below the free floor's 94.5 pct -- the model already loses on the headline, so the thing worth alarming on is the ONE column it wins: casefile_ground_caught_pct falling toward the floor's 64.7 pct, at which point paying for a model buys nothing on this corpus.
How tight can the band be? There is no scoring threshold to tune -- the grader is exact match on strings, dates and whole days, so nothing here has a tolerance. The one number that IS a threshold lives in the corpus, not the grader: MP-6.1's equivalent-document allowance, which the packet prints and the arm must apply.
Cadence: Once per corpus version. Every arm here ran ONCE and no arm is re-fired after its misses have been read -- the published number is the one the arm produced the first time, with the diagnosis attached.
The decisionWhen to reach for it
Use it
Always. Every published figure in this lens rests on it.
Do not use it
It cannot tell you a finding was defensible-but-different. A criterion the key calls DURATION_NOT_MET and the arm calls NO_SUPPORTING_DOCUMENT is scored wrong on the ground even though both agree the criterion is unmet -- and the sentence the practice receives really would be wrong, which is why it is scored that way.
Where the answer had to be read from -- the printed tables, or the correspondence
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineWhere the answer had to be read from -- the printed tables, or the correspondence
whether an arm is reading the criteria table or the prose. Splits every ground into the 84 the printed tables prove and the 34 only the clinical correspondence settles. This is the pair that separates the model from the free floor, and the ablation is what proves the split is real rather than a coincidence of difficulty.
$0.00per 1,000 prior-authorisation packets
nodata leaves your network
yessame answer every time
MethodHow the test was run
python3 -m evals.run --run-id <id> [--floor <mode>|--blind]; evals/scoring.py then compares strings, dates and integers against data/gold.jsonl. No model is in the grading path.
Every grader on these pages scored the same 40 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
CR-0014-05
CR-0014-05
doc id
PA-0014
case
superseded
printed tables
every figure on this criterion's row checks out: SD-0014-05 is filed, dated 2026-03-15, inside the 12-month lookback from 2026-04-17, and is an Imaging report -- exactly the kind the criterion names
clinical correspondence
Addendum AD-0014-05: the finding recorded in SD-0014-05 is withdrawn. On review the study reported there was the member's prior study and does not support this request.
free floor
CRITERION_MET -- CR-0014-05 is met by the document filed against it.
model
CRITERION_UNEVIDENCED / SUPERSEDED_BY_ADDENDUM, clause MP-3.1, evidence SD-0014-05 and AD-0014-05, RETURN -- AD-0014-05 withdraws SD-0014-05, so no imaging report currently confirms the level treated.
The whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
pass
The key scores this criterion CRITERION_UNEVIDENCED, ground SUPERSEDED_BY_ADDENDUM, clause MP-3.1, return RETURN. The paid arm answered all four, so the criterion counts on both denominators — the finding and, separately, the return call. The free floor answered CRITERION_MET and loses the row: it reads the criteria table, where SD-0014-05 is filed, correctly dated and of the required kind, and nothing in that table records that a later addendum withdrew it. This is one of the 34 grounds only the correspondence settles.
Where the answer had to be read from -- the printed tables, or the correspondence
pass
Scored as a CASEFILE ground, not a table ground, because the addendum is the only thing that decides it — every figure the printed tables carry on this row checks out. It sits in the 34-row denominator where the arm scores 91.2 pct against the floor's 64.7, and the correspondence-blind ablation answers CRITERION_MET here, which is what the 38.2 pct is made of.
Citations that do not appear in the packet
pass
Both identifiers the arm cited — SD-0014-05 and AD-0014-05 — are printed in packet PA-0014, and MP-3.1 is a clause the packet's own policy section carries. Nothing was invented. This grader's right answer is exactly zero and the scored run returned 0 fabricated identifiers across all 220 criteria.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
full packet
scored 91.2%
the ablation
scored 38.2%
In operationWhat to monitor
Reference standard: The channel of every criterion is fixed by tools/build_corpus.py when the packet is generated, before any arm sees it.
No true/false rates for this grader. It records 3 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
the gap between table_ground_caught_pct and casefile_ground_caught_pct -- that gap IS the argument for paying for a model
the ablation's casefile figure, which is what proves the prose was read rather than guessed around
Alarm on
the ablation's casefile_ground_caught_pct rising toward the full arm's. If removing the correspondence stops costing anything, the arm was never reading it and the kit's central claim is gone.
How tight can the band be? No threshold. Channel is a property of the corpus, assigned at generation time, not a score with a cut-off.
Cadence: Once per corpus version, alongside the ablation it exists to interpret.
The decisionWhen to reach for it
Use it
Always. Every published figure in this lens rests on it.
Do not use it
It says WHERE an answer had to be read from, not whether reading it was hard. A table-settled ground can still be missed for an unrelated reason, and this grader will report that as a table failure.
Check what a prior-authorization packet actually proves
PresenterOpens the private repo. Visible to admins only.
In one lineCitations that do not appear in the packet
whether any clause or document identifier the arm wrote is absent from the packet in front of it. Read from the PACKET, not the key, so the same check would convict the answer key if the key ever cited something a packet does not print.
$0.00per 1,000 prior-authorisation packets
nodata leaves your network
yessame answer every time
MethodHow the test was run
python3 -m evals.run --run-id <id> [--floor <mode>|--blind]; evals/scoring.py then compares strings, dates and integers against data/gold.jsonl. No model is in the grading path.
Every grader on these pages scored the same 40 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
CR-0014-05
CR-0014-05
doc id
PA-0014
case
superseded
printed tables
every figure on this criterion's row checks out: SD-0014-05 is filed, dated 2026-03-15, inside the 12-month lookback from 2026-04-17, and is an Imaging report -- exactly the kind the criterion names
clinical correspondence
Addendum AD-0014-05: the finding recorded in SD-0014-05 is withdrawn. On review the study reported there was the member's prior study and does not support this request.
free floor
CRITERION_MET -- CR-0014-05 is met by the document filed against it.
model
CRITERION_UNEVIDENCED / SUPERSEDED_BY_ADDENDUM, clause MP-3.1, evidence SD-0014-05 and AD-0014-05, RETURN -- AD-0014-05 withdraws SD-0014-05, so no imaging report currently confirms the level treated.
The whole finding on each criterion -- disposition, ground, the clause identifier, the documents -- AND, on its own denominators, the return call
pass
The key scores this criterion CRITERION_UNEVIDENCED, ground SUPERSEDED_BY_ADDENDUM, clause MP-3.1, return RETURN. The paid arm answered all four, so the criterion counts on both denominators — the finding and, separately, the return call. The free floor answered CRITERION_MET and loses the row: it reads the criteria table, where SD-0014-05 is filed, correctly dated and of the required kind, and nothing in that table records that a later addendum withdrew it. This is one of the 34 grounds only the correspondence settles.
Where the answer had to be read from -- the printed tables, or the correspondence
pass
Scored as a CASEFILE ground, not a table ground, because the addendum is the only thing that decides it — every figure the printed tables carry on this row checks out. It sits in the 34-row denominator where the arm scores 91.2 pct against the floor's 64.7, and the correspondence-blind ablation answers CRITERION_MET here, which is what the 38.2 pct is made of.
Citations that do not appear in the packet
pass
Both identifiers the arm cited — SD-0014-05 and AD-0014-05 — are printed in packet PA-0014, and MP-3.1 is a clause the packet's own policy section carries. Nothing was invented. This grader's right answer is exactly zero and the scored run returned 0 fabricated identifiers across all 220 criteria.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 0.0%
the ablation
scored 0.0%
In operationWhat to monitor
Reference standard: Read from the PACKET by src/packet.py::all_ids, never from the key -- so the same check would convict the answer key if the key ever cited something a packet does not print.
No true/false rates for this grader. It records 2 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
any non-zero invented_citation_pct at all -- this is the one column with a right answer of exactly zero
Alarm on
a single invented identifier. An invented citation is worse than none, because it is the first thing the requesting practice will check, and it is the failure that would end trust in the whole output.
How tight can the band be? No tolerance and no threshold: an identifier is either printed in the packet or it is not.
Cadence: Every run, on every arm, including the free floors -- which cite only what they read and so are structurally incapable of failing it.
The decisionWhen to reach for it
Use it
Always. Every published figure in this lens rests on it.
Do not use it
It checks that an identifier EXISTS in the packet, not that it is the right one. Citing a real document that does not support the finding passes this grader and is caught by the document_attached column instead.
A living map of modern AI — kept current every morning