The business caseThe problem this solves
A meter reader walks a route, photographs each meter with a capture app that burns the date, the technician and the premise id into the frame, and the photographs come back to a desk. Somebody keys twelve values off each one — serial, register reading, wheel count, unit, multiplier, meter type, the three overlay fields, the seal, whatever is on the glass — and then checks each against the premise's own record to decide whether the read can go to billing or has to go to a meter data specialist, field services or revenue protection. The desk keying of a field meter photograph: twelve values off the plate and the camera overlay, and the by-eye check of each against the premise record that decides which reads a specialist has to see.
Audience
A metering product manager deciding whether an OCR reader plus a model belongs in front of field read capture. The honest answer on this corpus is a heavily qualified no for the MODEL half and a clear yes for the question the kit was built to ask: the free regex floor over the same OCR text scores 0.6222 against the paid call's 0.6241 on one track and BEATS it 0.6259 to 0.6204 on the other, so the model call is not where the money belongs. Where it belongs is stage 1 — and three of the twelve fields cannot be bought at all under this arrangement. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual meter photographs
The corpus is 45 meter photographs, 22.27 MB (json 45 · png 45). A meter faceplate is the cleanest place to ask the question this kit exists to ask, because it carries three genuinely different KINDS of field at once. Nine are printed characters a perfect reader recovers. One, the seal, is printed in ONE DIRECTION ONLY — the lead disc is stamped SEAL and the stamp is only there while the seal is whole, so absence is ambiguous by physics rather than by rendering. Two are properties of the photograph itself: glare, dirt, condensation, vegetation, a fogged or cracked or sun-faded register. A corpus of clean printed forms cannot separate 'the model got it wrong' from 'nothing in the text could have got it right', and this one can. It is generated rather than collected for the same reason the answer key can be trusted at all: a real field photograph would have to be hand-transcribed to be scorable, and a hand transcription of a glare-blown register is an opinion arrived at by exactly the guessing the kit is measuring. The field conditions are DRAWN, never captioned — no image anywhere carries the word 'glare' — and the builder crops the register window out of each FINISHED photograph and requires its luminance spread to clear 26.0 (measured minimum 27.34) so that a condition degrades a read without erasing it.
The corpus
- The 45 meter photographsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your meter photographs. That is the whole change — there is no database to migrate.
{
"premise_id": "PR-76161",
"meter_of_record": {
"serial": "M3-3915889",
"unit": "m3",
"register_digits": 8,
"multiplier": "x1",
"type": "water"
},
"previous_read": {
"value": "75770154",
"date": "2026-08-12"
},
"trailing_avg_daily": 1.3,
"cycle_window": {
"start": "2026-09-11",
"end": "2026-09-17"
}
}
The outcomeWhat a good result looks like
A proposed read with the twelve cells the photograph actually supports, and the exception set pure code derives from them. In this kit's own units: 540 graded cells over 45 meter photographs, and 47 true exception instances across 9 codes, of which 7 cases carry none. The output is a PROPOSED read and a queue; nothing in this kit posts to a bill and nothing overwrites the premise's read of record.
And when it cannot
Mostly a REFUSAL rather than a wrong value, which is the safe direction and is still a miss. On the shipped track 200 of 540 cells came back null (0.3704) against 5 cells returned non-null and wrong (0.0093). The dangerous shape is the exception nobody sees: exception recall on the two purchasable tracks is 0.3617, so 30 of the 47 true exception instances never reach a specialist. Two of the nine codes are at ZERO recall on every arm ever run here — X6 OBSTRUCTED_READ (0 of 8) and X7 TAMPER_SEAL (0 of 5) — and that is a ceiling of the arrangement, not a gap in the model.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Field meter reads on machine-printed faceplates, where the twelve values you need are all PRINTED — src/floor.py — the free regex over a purchased OCR reading, no model call at all
on the batch reading the floor scores 0.6259 against the paid call's 0.6204 and wins outright; on the synchronous reading it is one cell behind (0.6222 against 0.6241) with p = 1.000000. It raises a better exception queue on both (0.5882 F1 against 0.5000) and it asserts 1 wrong non-null value where the model asserts 6. It costs the same page bill and nothing else. - The same job, but the plate carries manufacturer language a fixed table cannot enumerate — the model call after the reading, and keep the rules in pure code
meter_type is the one field where the paid call is significantly ahead — 0.9333 against 0.7333, p = 0.003906 over 9 discordant pairs. The plate prints DIAPHRAGM GAS INDEX or POSITIVE DISPLACEMENT or a 240V 3W rating rather than the word gas, and the model reads the sense of that where the regex needs a marker list - You need the tamper exception — a broken seal must reach revenue protection — a second signal that is not the photograph's text
X7 TAMPER_SEAL recall is 0 of 5 on EVERY arm ever run here, including the arm handed perfect text. The lead disc is stamped SEAL and the stamp is only there to be photographed while the seal is whole, so absent text means broken OR not visible and nothing separates them. A transcript can verify a seal is intact and can never verify one is broken. On real OCR it is worse: the stamp appears in 36 of 45 perfect texts and 0 of 45 on either purchasable track - You need the obstructed-read exception, or any judgement about the state of the glass — an image model that sees the pixels — arrangement A2, which this kit did not build
obstruction and register_condition are properties of the IMAGE and every arm scores 0 of 90 on them. X6 OBSTRUCTED_READ recall is 0 of 8 on every arm including the oracle. This is not a hard field, it is an absent one — the schema keeps both because removing them would publish a field_exact_match that flatters the arrangement - Choosing between the two Mistral OCR service levels on price — the batch tier
$2 per 1,000 pages against $4, same model id, same 45 pages. The batch tier's own span CER is slightly LOWER (0.1269 against 0.1278) and its field score is 2 cells lower — and 42 of the 45 stage-1 texts are byte-identical between them, so most of that difference is the model disagreeing with itself - Choosing an OCR VENDOR for meter faceplates — measure it yourself; this kit cannot tell you
both priced tracks are ONE product, mistral-ocr-4-1, at two service levels. The approved pick board named AWS Textract Detect Text and Google Document AI Enterprise OCR; both were credential-blocked at build time and neither ran a single page. The runners exist and are finished — docs/INFRA.md carries the probe evidence and the unblock table — so this is an account fact rather than a code fault
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus at your own photographs and write a data/gold.jsonl of the same shape: one line per case carrying image, context and the twelve fields, plus a context file per premise holding meter_of_record, previous_read, trailing_avg_daily and cycle_window. The reading numbers do not travel and the answer key cannot come with you. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not carry this across to plates the floor's caption table has never seen. Its strength is a fixed printed vocabulary — NO., BADGE NO., SER., FIG, MULTIPLIER — and this corpus has three plate families. It also derives meter_type from the unit where the plate does not print it, which is a declared shortcut, and it refuses that derivation on m3 because cubic metres are billed by gas and water alike. That is the case against the best-fitting scenario (“Field meter reads on machine-printed faceplates, where the twelve values you need are all PRINTED”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A condition that covers the register rather than dimming it. On MR-0013 the vegetation sits INSIDE the register window and both purchasable tracks returned four image placeholders where the digits are — no reading at all, so no rollover check and no consumption check. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THIS IS NOT A CROSS-VENDOR COMPARISON, AND IT WAS MEANT TO BE — the single largest limitation in this kit. Both priced tracks are ONE product, mistral-ocr-4-1, at two service levels: $4 per 1,000 pages synchronously and $2 in batch. 14 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r003-meter-read-capture-mistral. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on this checkout with the machine's own python3, no virtualenv, no install and no credential configured: python3 -m src.rules --self-test passes, python3 -m src.floor --self-test passes, python3 -m src.rules --against-gold reports 45 cases and 0 mismatches, and a full six-arm re-score through evals/score.py returns every figure in this spec in under a second. All three write nothing. What could NOT be exercised on this checkout is python3 tools/build_corpus.py --check, which re-derives the answer key from the seed: Pillow is not installed here and the builder is the one file in the kit that imports it, so the byte-identical re-derivation is a claim this spec does not make. The two purchasable stage-1 tracks need a vendor key and the model stage needs a provider; every arm already on disk re-scores with neither.



