The business caseThe problem this solves
A telehealth visit is documented in a note and billed on a claim line, and the two are produced by different people minutes apart. The note says how the encounter was conducted and where the patient was; the line says it with a two-character modifier and a two-digit place of service. A coder holding a queue of them opens the note, the plan's rule card and the claim in three windows and reads two facts out of the first to check the third. Reading a visit note for two facts and joining them to a plan's telehealth rule card, one claim line at a time.
Audience
A certified coder working a telehealth queue, and the coding manager deciding whether a check like this is worth putting in front of them. The answer this kit gives is a qualified yes on one field and a plain no on two of the six. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual telehealth claim-line reviews
The corpus is 62 telehealth claim-line reviews, 0.05 MB (txt 62). It is generated because it has to be. A real telehealth visit note is the most identifiable document in the encounter, and the two facts this job needs -- how the encounter was conducted and where the patient was -- do not require any of the rest of it. Generating the corpus is what makes that separation demonstrable instead of asserted, and it is what lets the answer key be DERIVED from the same structure the review is built from rather than typed by somebody reading it afterwards.
The corpus
- The 62 telehealth claim-line reviewsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere -- every one of the 62 claim-line reviews, the coded-line register and the whole answer key are generated in-process. Re-running the generator rebuilds all 62 byte for byte under any PYTHONHASHSEED.
Swap this folder for your own material and the kit is pointed at your telehealth claim-line reviews. That is the whole change — there is no database to migrate.
TELEHEALTH CLAIM LINE REVIEW TCL-0001
ENCOUNTER HEADER
Encounter ENC-263476
Date of service 2026-05-16
Clinician role Nurse practitioner
Modality audio-video
Patient location home
VISIT NOTE
The patient made contact about a change in symptoms since the last review.
The patient was on camera for the entire encounter and I was able to inspect the site visually.
The patient took part from home, in their own living room.
No change to the current plan; review at the next scheduled appointment.
CLAIM LINE AS CODED
Line Service Service group Description Modifier(s) POS Units Charge
1 TSV-1043 evaluation Established patient assessment, low complexity 95 10 1 $160.00
CODING NOTES
Charge captured from the standard fee schedule in force on the date of service.
The outcomeWhat a good result looks like
One review in, six graded answers out: how the encounter was conducted, where the patient was, the modifier and the place of service TPR-2026 requires for that, one verdict on the line as coded, and the note line establishing the modality quoted verbatim so the coder can see the evidence without opening the note.
And when it cannot
And what it does when it cannot. On five of the 62 lines both paid runs answered a patient location the key does not hold; on every one of those the verdict was still right, and the cause is a defect in this kit's own corpus generator rather than in the reading. Where the note documents nothing it answers NOTE-SILENT and quotes no line, which is a query to the clinician rather than a correction to the claim.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your notes carry a reliable structured modality field and you trust it — the free header floor, and no model at all
It is 40 of 62 here only because 8 reviews are built to contradict the header. Where the header is trustworthy the join is exact, free and instant. - You have narrative notes and no reliable structured field — the free rules floor first, then measure whether a call beats it
45 of 62 for $0.00, and it is the arm this kit publishes as its floor. On a corpus where the modality is stated plainly, phrase matching is most of the job. - Your notes describe what happened rather than declaring it, and a mismatched line reaching the payer is expensive — the paid call, with the rule card re-applied in code afterwards
0 of 43 mismatched lines passed on both scored runs, against 8 for the best free floor, and 8 of 8 on the header/narrative conflicts against 7 and 0. - You want the verdict auditable rather than merely accurate — either arm, plus src/recheck.py
The station discards whatever verdict came back and re-derives it from the reading under the card's six ordered rules, so the verdict is always the card's and never an opinion. It fired 0 times here, which means the model applied the card correctly every time it read correctly.
At a glanceHow the whole thing runs
Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own reviews and data/lines.json with your own coded lines, then rewrite data/policy.json and data/policy.md with your plan's card -- the modifier table, the place-of-service table, the excluded service groups and the order the verdict rules fire in. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it where the header is a scheduling template value. It gets 0 of the 8 conflicts and every one of those is a mismatched line passed. That is the case against the best-fitting scenario (“Your notes carry a reliable structured modality field and you trust it”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A note longer than a page. The review is sent whole with no reduction step, so a real encounter note with the modality mentioned once in six paragraphs is a different problem and this kit has no answer to it. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A third run, a different temperature, or any reasoning setting other than the tier's default. Two identical runs agreeing on 62 of 62 says the sampling is not moving this reading; it says nothing about a different prompt or a different day. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-03 — r001-modifier-pos. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9297 -- all 62 reviews, the answer key, all three free floors computed live in the browser, the rule-card derivation, the verdict vocabulary and the full corpus table -- and replays both committed paid runs off their result files. The only control that needs a key is disabled and says why. Nothing is installed: requirements.txt names no package because nothing under src/ or evals/ imports one.




