The business caseThe problem this solves
A claimant films a guided walkaround of a damaged vehicle and the frames carry printed text an adjuster needs: the VIN plate, the licence plate, the odometer and a repair-shop damage tag on each panel. Today an intake handler scrubs the video, pauses on the readable frames and keys twelve fields per claim by hand, then checks the VIN check digit and the odometer against the policy record in their head. The frame-by-frame desk read of a walkaround video: keying the claim-intake record and deciding, by eye, which claims need a human look.
Audience
A claims-intake product manager deciding whether a reader belongs in front of FNOL video. The honest answer on this corpus is a qualified no on the part that costs money: the free regex floor matched the paid model call cell for cell. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual vehicle walkarounds
The corpus is 30 vehicle walkarounds, 13.15 MB (jpg 225 · json 1 · mp4 30 · txt 225). A walkaround is not a document, and that is the point of building this one. The twelve fields sit on different physical objects — a door-jamb label, a bumper plate, bright digits behind dark glass, a card wired to each panel — so a field can be missing because nobody filmed it, a failure mode a flat single-page corpus cannot produce. Generated rather than collected, so the truth is what was PRINTED before it was printed: a real walkaround video would need every frame hand-transcribed to be scorable, and a hand transcription is an opinion that then gets graded as fact. Each frame ships with a .txt sidecar of the literal text drawn on it, in reading order and never rebuilt from the field values, which is what lets an oracle arm read perfect text and separate the reader's errors from the model's. The .mp4s are a convenience for a human watching — the unit is the frame, and the frames stand alone
The corpus
- The 30 vehicle walkaroundsgenerated from a fixed seed, so no real record, person or institution appears in it.
Swap this folder for your own material and the kit is pointed at your vehicle walkarounds. That is the whole change — there is no database to migrate.
NORTHWIND MUTUAL
FIRST NOTICE OF LOSS — INTAKE SLIP
POLICY NO PA-5507-25097
CLAIM REF FNOL-4400
LOSS DATE 2026-03-21
SHOP OAK HOLLOW COLLISION
SYNTHETIC — GENERATED FOR EVALUATION
FRAME 00/07 00:00 FNOL-4400
The outcomeWhat a good result looks like
Every printed cell filled and every claim that deserves a second look flagged. In this kit's own units: 822 graded cells across 30 walkarounds, 222 frames and 107 damage tags, with all 31 true exceptions raised. The oracle arm — perfect text, no reader — reaches 822 of 822; Mistral OCR 4.1 sync reaches 818 and the batch tier 817. The output is a claim-intake record plus a queue a human decides, each exception naming the rule that fired and the two values it compared. It never approves, pays or closes a claim.
And when it cannot
A field keyed from a misread frame. On this run every wrong field came from the reading and none from the model — propagation is 1.0 on both paid arms (4 wrong cells of 822 sync, 5 batch), so a better reader buys more here than a better model. The failure it can still have is an exception queue nobody trusts: recall is 1.0 on every arm and no true exception was missed, but the sync arm raised 34 to find 31 and the batch arm 35 — 3 false alarms sync (one each on VIN checksum, VIN-to-policy and plate-to-policy) and 4 batch (two on VIN checksum). Every one of them is a rule firing correctly on a misread VIN or plate, which is what a false alarm looks like when the rules are pure code and the reading is not.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Guided walkaround video of a damaged vehicle, where the frames carry machine-printed plates and repair tags — the batch reading tier, with pure code deciding the exceptions
measured at 0.9939 of cells and 0.9394 exception F1 with recall 1.0000 on all seven codes, about five times faster than the synchronous tier, and every one of its five wrong cells traced to the reading rather than the model - The same job, but you want the fewest false alarms on the human queue — the synchronous reading tier
3 false alarms against the batch tier's 4, on 34 raises against 35 — one misread frame's worth of difference, which is a real but very thin reason - A fixed, machine-printed tag layout that does not change between shops — src/floor.py — the free regex, no model call at all
on this corpus it matched the model extractor on all 822 cells, on every arm, with the same wrong cells and the same exception sets, at zero tokens and zero calls - Handwritten damage notes, or frames with no printed tag at all — measure it before trusting any of this
every frame in the corpus is machine-printed then degraded; handwriting is a different measurement - Any deployment where a misread must not be allowed to post itself — the two-stage build with the policy record withheld from the model
the VIN checksum is pure arithmetic the model never sees and cannot be talked out of; all three VIN misreads across both tiers fired vin_checksum_invalid and reached a human named, with both compared values attached
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-10. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus with your own frames and a gold.json of the same shape: walkarounds[].frames in filming order, walkarounds[].claim for the twelve fields, and walkarounds[].policy carrying policy_vin, policy_plate, last_odometer and last_service_date — the policy record the model is never shown. The reading numbers do not travel. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not infer that the batch tier is free of risk because it is cheaper and faster — it is the tier that produced the extra VIN misread here. That is the case against the best-fitting scenario (“Guided walkaround video of a damaged vehicle, where the frames carry machine-printed plates and repair tags”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A reader that returns the letter O where the frame prints a zero. Both paid arms did it — 0EM19WRS7SF341456 came back OEM19WRS7SF341456, and the batch arm did it again on 0NHHVX9ZXSJ360964 — and because the VIN alphabet excludes O by construction, one wrong character false-fires vin_checksum_invalid and vin_policy_mismatch together. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THIS IS NOT A CROSS-VENDOR COMPARISON, and it was meant to be. Both priced tracks are ONE product, mistral-ocr-4-1, at two service levels. 15 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is openai-compatible; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-10 — r001-fnol-walkaround-mistral-ocr. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on this checkout, with the machine's own python3 — no virtualenv, no install and no credential configured: the five offline self-tests all pass and write nothing to disk (src/read.py, which re-derives its SigV4 signing chain and an RS256 signature against openssl-measured vectors with no key; src/rules.py; src/floor.py; evals/score.py; evals/injection.py). Every module under src/ and evals/ imports only the standard library; Pillow appears in tools/build_corpus.py alone, and that file is the only one with a dependency at all. Two of the five reading engines could not be exercised: AWS Textract answered AccessDeniedException on textract:DetectDocumentText with live, correctly signed credentials, and Google Document AI answered SERVICE_DISABLED after accepting the token — both account facts above the code, recorded rather than worked around.

