Check each production batch's lab certificate against its limits
Every batch comes with a lab certificate, and the analyst's note does not always agree with the numbers. This app reads each certificate, lays out the result and the limits, and says whether the batch meets its specification.
PresenterOpens the private repo. Visible to admins only.
For quality controlCross-domain · Manufacturing & CPG
Why it matters
Today's manual process, and the same job with the app
Quality-control staff at a food or materials maker, reviewing lab certificates before a batch moves, or at the customer receiving it.
✕Today's manual process
1Open each certificate and find the measured result among the analyst's comments.
2Find both limits in the specification section, and notice when one is missing.
3Compare them manually, past a note that may sound calmer or more worried than the numbers.
4One misread lets an out-of-spec batch move into finished goods.
Every certificate compared manually
✓With the app
1Each certificate is read into ten named fields, each tied to the section it came from.
2Both limits are laid out, and a missing one is marked not stated.
3The comparison uses the numbers, never the analyst's note.
4An out-of-spec batch shows a plain no, however calm the note. Your team still makes the release call.
People check the answer, not the arithmetic
See it work
One real case, read by the app, step by step
Cocoa powder batch B-5515 measures 1025 against a 1000 limit, and the analyst's note says released.
Check each production batch's lab certificate against its limitsReference appBuilt to be shaped to your process
5
1The certificate batch B-5515, Brindle Alkalised Cocoa Powder, tested for total plate count.
2The measured result 1025 CFU/g, tied to the certificate section it came from.
3The limits no lower limit is stated, and the upper limit is 1000.
4The analyst's note says within normal range and released. It does not decide the answer.
5The answer: no 1025 is over 1000. Your team makes the release call.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check each production batch's lab certificate against its limits
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Deciding whether a manufactured batch met its specification means comparing one measured laboratory result against the stated limits — arithmetic — while the certificate also carries the analyst's own free-text disposition note, which is an opinion and can point the other way. Someone reading each certificate of analysis, finding the measured result and the two specification limits among the analyst's own commentary, and doing the comparison by hand before the batch moves — including on the certificates whose disposition note reads more reassuring, or more worried, than the numbers warrant.
Audience
Quality-control and quality-assurance teams who review certificates of analysis before a batch moves, at a manufacturer or at the customer receiving the material, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual certificates
The corpus is 55 certificates, 0.03 MB (txt 55). Plain text, one format, invented rather than fetched — a real certificate of analysis cannot be published either (see SOURCES.md). Ten fields are chosen because the same ones decide a release review: which batch, which material, which parameter, the result and its unit, the two limits — either of which may be absent on a one-sided specification, and 31 of these 55 records carry one — the analyst's own note, and the verdict this kit exists to test.
The corpus
The 55 certificatesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your certificates. That is the whole change — there is no database to migrate.
One certificate, as the model receives itCOA-0001.txt · 1 of 55
Batch
-----
BATCH-E-5330
Product
-------
Brindle Alkalised Cocoa Powder
Test Parameter
--------------
bulk density
Measured Result
---------------
0.61 g/mL
Specification Lower Limit
-------------------------
0.55 g/mL
Specification Upper Limit
-------------------------
0.75 g/mL
Test Date
---------
2026-06-24
Analyst
-------
Casey D. Rossi
Analyst Disposition Note
------------------------
Result consistent with the last twelve batches of this grade. No concerns; pass.
The outcomeWhat a good result looks like
A ten-field extracted record per certificate, plus one pure-code self-consistency check: the same comparison re-run over the model's own extracted numbers, flagged when it disagrees with the model's own verdict. Informational only — nothing here releases or rejects a batch.
And when it cannot
This run found one field-level miss across both tiers and 1,100 combined cells — a product-name misread on a field the comparison never reads — and zero conformance-verdict errors on either tier. What the run did not test: a unit mismatch between the result and the limit, a specification that changed between the test date and the review, a re-test superseding the first, or a measured value whose analytical uncertainty is wider than its distance to the limit. See Business.not_good_enough.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Screening a stack of certificates for batches that are out of specification before a quality reviewer opens them — either tier — they tie at 55 of 55 verdicts, 1.00 recall and precision measured here 100% conformance accuracy on both tiers against the free analyst-tone floor's 56.4% (24 wrong, including 10 out-of-specification batches called conforming).
Deciding the two tiers on cost, speed or field-level precision — the fast tier About 9% cheaper per certificate ($0.0012524 vs $0.0013675), 45% lower p50 latency (2,415 ms vs 4,396 ms), and identical on every published verdict figure.
At a glanceHow the whole thing runs
99.8%extraction accuracy
2,415 msp50, end to end
$1.25per 1,000 certificates · Google Gemini 3 Flash
Run once, for real, on 2026-08-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check each production batch's lab certificate against its limits14 steps · 4 questions · run once, for real · 2026-08-21
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per certificate. Corpus lens →
When is this the wrong choice?
Avoid: The analyst-tone floor for the conformance verdict specifically — reading the disposition note is exactly what the planted ambiguity is built to defeat. That is the case against the best-fitting scenario (“Screening a stack of certificates for batches that are out of specification before a quality reviewer opens them”). 2 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Scanned or image-only certificates — there is no OCR step. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether either tier would still score 55 of 55 on a larger or adversarially-constructed set of register-mismatched certificates. This run's 24 mismatched cases (of 55) all resolved correctly on both tiers, but 24 cases is not enough to rule out a harder confusion this corpus did not think to plant. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the deliberating tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-21 — r001-coa-conformance. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Verified from a cold clone with no key configured: 55 certificates segment into 495 sections in 0.002 seconds, tools/build_corpus.py regenerates the whole corpus byte-identically in 0.038 seconds, and the assembled prompt for COA-0033 replays with the same three part sizes run r001 recorded.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
2,415 msp50, end to end
3,225 msp95
2 minclone to first result
What the clock covers. model call only, one per certificate
Current processWhat it replaces
Someone reading each certificate of analysis, finding the measured result and the two specification limits among the analyst's own commentary, and doing the comparison by hand before the batch moves — including on the certificates whose disposition note reads more reassuring, or more worried, than the numbers warrant.
Where it is not good enough
One field-level miss across both tiers and 1,100 combined cells: the fast tier read a product name as "Halcotte" instead of "Halvette" on COA-0038 — a field the conformance comparison never touches. Both tiers scored 55 of 55 conformance verdicts with 1.00 recall and precision, and neither tier's self-consistency guardrail ever fired. That is a clean result on 55 records and a small sample for a comparison this consequential. It also says nothing about the harder cases this corpus does not build: a result reported in a different unit from the limit, a limit expressed as a percentage of a target, a specification that changed between test and review, or a re-test that supersedes the first. And the guardrail checks the model against itself, not against the certificate — a reply that misreads the measured value and judges that misreading correctly is self-consistent and stays unflagged.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The guardrail is a self-consistency check: it re-runs the comparison over the numbers the model itself returned and fires when they do not support the model's own verdict. It needs no gold, so it computes on unlabelled certificates. Both tiers were self-consistent on all 55. The free analyst-tone floor (evals/baseline.py, which reads the disposition note and never compares the numbers) got all 495 structured cells right and 24 of 55 verdicts wrong — and the same guardrail caught every one of those 24 with no gold at all. No red-team run exists for this kit; this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
SECTION_HINTS
src/select.py
map fields to your own certificate's headings; unmatched falls back to the whole document
PROVIDERS
src/adapters/__init__.py
any OpenAI-compatible host, or Anthropic's Messages API
recompute_conformance
src/extract.py
the comparison itself — the boundary convention, or a tolerance band if your programme allows one. It is stated once and used by the corpus generator, the prompt and the guardrail, so changing it here changes all three together
the field schema
data/fields.json
a different set of fields entirely, with its own types and allowed values
Components
Component
File
Role
segment
src/segment.py
cut the certificate into addressable sections, pure code
select
src/select.py
pick which sections carry each field, pure code — conformance is mapped to the measured result and the two limits, never to the disposition note
prompt
src/prompt.py
assemble one call for all ten fields, with the comparison rule and its boundary-inclusive convention stated in full
extract
src/extract.py
the AI layer, one provider one key — plus the pure-code self-consistency check downstream: re-run the comparison over the model's OWN extracted numbers and flag any disagreement with the model's own verdict
judge
evals/judge.py
score field accuracy and the conformance confusion matrix separately, pure code
Where it breaks at scale
One call per certificate, no concurrency and nothing shared between calls: 55 certificates took 133 seconds wall clock on the fast tier. A release queue of thousands needs batching and a rate-limit strategy this kit does not have. It also reads one certificate in isolation — a real quality review compares a batch against its own re-test and against its neighbours on the same production run, and nothing here does that.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Before anything is asked of the model. Ten named fields with their own types and allowed values, plus a second panel for the three figures computed afterwards in pure code — the model's verdict, the recomputation, and whether they agree.successOpen full size →COA-0033, extracted live. A one-sided specification (no lower limit, upper limit 1000 CFU/g), a measured 1025, and an analyst note reading "Within normal range for this product line. Released to finished-goods inventory." The model answered conforms_to_spec: no — the arithmetic, not the prose — and the recomputation over its own numbers agrees, so the self-consistency check stays quiet. The free analyst-tone floor calls this batch conforming.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same button with no API_KEY configured. A calm 200 and a plain sentence, not a stack trace: nothing was called, nothing was spent, and the field table stays browsable.failureOpen full size →
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
55certificates
0.03 MiBtxt 55
495sections · p50 44 chars
$0.00setup · 0.002s
How it is cutWhat one section is
cut on underlined section headings; a certificate with none falls back to one whole-document segment so a span still resolves
SetupWhat the setup figure measured
There is no index. Preparation is segmentation only — 55 certificates cut into 495 sections by src/segment.py, pure code, no model and no key.
LicenceLicence
MIT — this repository's own licence. Every product, batch, analyst and specification limit is invented; a real certificate of analysis cannot be published either — it names a real manufacturer, a real lot and a real release decision.
Bring your ownBring your own certificates
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per certificate. SECTION_HINTS in src/select.py maps fields to headings and will need editing for a different certificate layout; when it does not match, selection falls back to the whole document — slower, more expensive, always correct. If your certificates are unlabelled, the field grade and the confusion matrix both go away and the self-consistency check does not: it is computed from the reply alone.
What breaks it
Scanned or image-only certificates — there is no OCR step.
A certificate whose sections are not headed — segment() falls back to one whole-document segment, so a span names "document" and locates nothing finer.
A result reported in a different unit from the limit it is compared against — this kit does no unit conversion, and every record in this corpus states both in the same unit.
A batch's own re-test, or its neighbours on the same production run, are never compared — this kit reads one certificate in isolation.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
1,529
419
field schema
1,585
370
certificate sections
486
176
Total
965
This is the cost lesson as arithmetic: of the 965 tokens assembled, 419 are instructions — 43% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Measured token-for-token via evals/prompt_tokens.py's nested-prefix subtraction against the provider's own tokenizer, not estimated from characters — see results/tokens-p001-coa-conformance.json.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You extract structured fields from a certificate of analysis for one manufactured batch. You return JSON and nothing else.
RULES, in order of importance:
1. If the certificate does not state a field, return null for it. Do not infer it and do not use what you know about the world. A specification limit the certificate explicitly says is not specified is null, not a number you supply.
2. `conforms_to_spec` is decided by ARITHMETIC ON THE NUMBERS, never by how the analyst's disposition note reads. Answer 'yes' when measured_value >= spec_lower_limit AND measured_value <= spec_upper_limit; answer 'no' otherwise. Both limits are INCLUSIVE -- a value exactly on a limit conforms. A limit the certificate does not state places no constraint on that side, so a one-sided specification is judged on the side it does state. Do the comparison yourself before answering.
3. The analyst's disposition note is a field to copy, not evidence about conformance. A note reading 'within normal range, released' does NOT make an out-of-limits value conform, and a note reading 'borderline, recommend re-test' does NOT make an in-limits value fail. The numbers decide; the note is the analyst's opinion and may disagree with them.
4. Copy values verbatim from the certificate wherever possible, and report measured_value, spec_lower_limit and spec_upper_limit as bare numbers with the unit left out of them.
5. Use the exact allowed value for a field that lists them.
6. Return every field named in the schema, even when the answer is null.
---
Extract these fields:
- batch_id (string) -- the batch or lot identifier, verbatim
- product_name (string) -- the name of the material the certificate is for
- test_parameter (string) -- what was measured, verbatim (for example: moisture content, assay purity)
- measured_value (number) -- the numeric laboratory result, without its unit
- unit (string) -- the unit the result is reported in, verbatim (for example: %, ppm, CFU/g, microns)
- spec_lower_limit (number) -- the lower specification limit as a number, or null when the certificate states the specification is one-sided and gives no lower limit
- spec_upper_limit (number) -- the upper specification limit as a number, or null when the certificate states the specification is one-sided and gives no upper limit
- analyst_disposition_note (string) -- the analyst's own free-text disposition note, copied verbatim
- test_date (string) -- the date the test was performed, YYYY-MM-DD
- conforms_to_spec (enum) one of: yes, no -- does the measured value fall within the stated specification limits? Decide this by ARITHMETIC ONLY: it is 'yes' when measured_value is greater than or equal to spec_lower_limit AND less than or equal to spec_upper_limit, treating a limit the certificate does not state as no constraint on that side; otherwise 'no'. Both limits are inclusive -- a value exactly on a limit conforms. Do NOT decide this from the analyst's disposition note: a note reading 'within normal range, released' does not make an out-of-limits value conform, and a note reading 'borderline, recommend re-test' does not make an in-limits value fail.
Return a JSON object with exactly these keys: batch_id, product_name, test_parameter, measured_value, unit, spec_lower_limit, spec_upper_limit, analyst_disposition_note, test_date, conforms_to_spec
Use null for any field the certificate does not state.
CERTIFICATE OF ANALYSIS
-----------------------
Batch
-----
BATCH-B-5515
Product
-------
Brindle Alkalised Cocoa Powder
Test Parameter
--------------
total plate count
Measured Result
---------------
1025 CFU/g
Specification Lower Limit
-------------------------
not specified (one-sided specification)
Specification Upper Limit
-------------------------
1000 CFU/g
Test Date
---------
2026-04-11
Analyst Disposition Note
------------------------
Within normal range for this product line. Released to finished-goods inventory.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"batch_id": "BATCH-B-5515",
"product_name": "Brindle Alkalised Cocoa Powder",
"test_parameter": "total plate count",
"measured_value": 1025,
"unit": "CFU/g",
"spec_lower_limit": null,
"spec_upper_limit": 1000,
"analyst_disposition_note": "Within normal range for this product line. Released to finished-goods inventory.",
"test_date": "2026-04-11",
"conforms_to_spec": "no"
}
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check each production batch's lab certificate against its limits — 55 certificates. Two tiers of one model family answered, and every answer was then graded Two different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
55certificates
55source documents
2model tiers
110graded answers
2grading methods
MeasurementsWhat was measured
COUNTED549 · 550 / 550extraction accuracy — cellsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED55 · 55 / 55conformance accuracy — certificatesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/judge.py) is pure code, exact match with light normalisation against a mechanically-derived gold — there is no judgement to validate, only comparison. What WAS validated: gold's conforms_to_spec is not a typed label at all, it is the arithmetic run over the same measured value and limits the certificate states, and evals/check_labels.py re-runs that comparison over every gold row before any run is allowed to spend, failing if a single label disagrees with its own numbers. tools/build_corpus.py's own _verify() pass separately confirms every gold value is stated verbatim in the document it labels and every test_date is a real calendar date.
256.95output tokens · the fast tier · 2,415 ms p50
295.33output tokens · the deliberating tier · 4,396 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.8× as long, and lands one row apart on 55. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One certificate
1,000 certificates
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.001252
$1.25
38%
Same work, 1× the bill
The same certificates, the same tokens — only the rate card changed. And on that card about 38% of what you pay is the prompt this pipeline sends, not the answer it writes.
which tier is called — the two tiers tie on every published verdict figure, so the lever buys nothing measurable here; the fast tier is both cheaper and faster, and the one field-level miss of the night landed on it.
Rates checked 2026-08-18. The provider that actually ran r001 and r002 publishes no rate card this repo commits, so nothing here is what was actually paid — the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page.
The gradersTwo ways to grade
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Per-field exact match, light normalisation Does the model's value for each field match gold, after trimming whitespace/punctuation and treating numbers within half a cent as equal — including the two limits, where a correctly-returned null on a one-sided specification is a hit and an invented bound is a miss?
$0.00
no
yes
the fast tier 99.8% · the deliberating tier 100.0%
Conformance confusion matrix Does the model's stated conforms_to_spec match the same comparison run over gold's own measured value and limits? Out of specification is the positive class, so recall answers "of every batch that really is out of specification, how many did the run call out of specification".
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 100.0%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, between the models and the analyst-tone floor, not at all between the two model tiers. Both tiers scored 55 of 55 conformance verdicts with 1.00 recall and precision and were self-consistent on every record; the deliberating tier hit every field cell and the fast tier missed exactly one (a product name, unrelated to the verdict). The floor is what separates cleanly: 56.4% verdict accuracy against 100%, on exactly the register-mismatched certificates this corpus was built to plant.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Screening a stack of certificates for batches that are out of specification before a quality reviewer opens them
either tier — they tie at 55 of 55 verdicts, 1.00 recall and precision measured here
100% conformance accuracy on both tiers against the free analyst-tone floor's 56.4% (24 wrong, including 10 out-of-specification batches called conforming).
the analyst-tone floor for the conformance verdict specifically — reading the disposition note is exactly what the planted ambiguity is built to defeat.
Deciding the two tiers on cost, speed or field-level precision
the fast tier
About 9% cheaper per certificate ($0.0012524 vs $0.0013675), 45% lower p50 latency (2,415 ms vs 4,396 ms), and identical on every published verdict figure.
assuming the fast tier is strictly better — the one field-level miss of the night (a product name) landed on the fast tier, not the deliberating one. It is the one thing the extra money bought, and it never touches the comparison.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
product-name-misread
A product-name misread, unrelated to the conformance comparison
1
COA-0038: gold product_name="Halvette Whey Protein Concentrate"; the fast tier returned "Halcotte Whey Protein Concentrate". The deliberating tier extracted it correctly. Neither the measured value, the limits, nor conforms_to_spec was affected, and the…
tone-floor-register-mismatch
The free analyst-tone floor's own failure mode, measured on the same corpus
24
evals/baseline.py decides conformance from the disposition note's wording and never compares the numbers. On the 24 certificates whose note points against the arithmetic it is wrong every time — 10 out-of-specification batches called conforming (COA-0002…
What we could NOT verify
Whether either tier would still score 55 of 55 on a larger or adversarially-constructed set of register-mismatched certificates. This run's 24 mismatched cases (of 55) all resolved correctly on both tiers, but 24 cases is not enough to rule out a harder confusion this corpus did not think to plant.
Whether the model can handle a unit mismatch between the measured result and the limit — every record here states both in the same unit by construction, and the kit does no conversion. A certificate reporting ppm against a limit stated in percent is unmeasured.
Whether the self-consistency guardrail would catch a model that misreads a number AND judges consistently against its own misreading — by construction it would not, and no run has produced that case to confirm the shape of the failure. It fired zero times on 110 replies.
How either tier performs on a real certificate of analysis (a multi-parameter certificate with a results table, a vendor-specific layout, footnoted methods) — this corpus is one parameter per certificate, plain text, with a single consistent layout by construction.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
963.07
256.95
2,415 ms
$0.001252
the deliberating tier
963.07
295.33
4,396 ms
$0.001368
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The analyst-tone floor (evals/baseline.py) and the scorer (evals/judge.py) are both pure code and cost $0.00 to run against either result set. The figure above is both runs' own token counts (the fast-tier and deliberating-tier runs, 55 certificates each) priced at Google Gemini 3 Flash's published rate — the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The fixed system prompt and field schema (789 of 965 tokens on the example call, 82%) outweigh the certificate sections sent (176 tokens) — the floor every call pays before a single number is read.
Output length: the model returns a full ten-key JSON record every call, including the analyst's disposition note copied back verbatim, whatever the certificate says.
Your volumeWhat it costs at your volume
Linear in certificates: each call is independent and self-contained, with no shared context or index to amortise. This run's 55 certificates cost about $0.069 projected onto Gemini 3 Flash's rate, so ten times the set is about $0.69 on the same rate and the same prompt — arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
52,969input tokens · this run
14,132output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 55 certificates of analysis extracted, each scored by pure code against a mechanically-derived gold set. This run (r001-coa-conformance, the fast tier) answered all 55 of 55 certificates with no truncation -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.028
$0.028
$0.50
2026-09-12
gemini-3-flash
Google
$0.069
$0.069
$1.25
2026-09-18
gemini-3-8-flash
Google
$0.093
$0.093
$1.69
2026-09-18
claude-haiku-4-5
Anthropic
$0.124
$0.124
$2.25
2026-09-12
llama-5
Meta
$0.126
$0.126
$2.30
2026-09-18
grok-4-5
xAI
$0.191
$0.191
$3.47
2026-09-18
grok-4-6
xAI
$0.191
$0.191
$3.47
2026-09-18
claude-sonnet-5
Anthropic
$0.247
$0.247
$4.50
2026-09-12
gemini-3-1-pro
Google
$0.276
$0.276
$5.01
2026-09-18
gpt-5-6-terra
OpenAI
$0.276
$0.276
$5.01
2026-09-12
gpt-5-6-sol
OpenAI
$0.495
$0.495
$8.99
2026-09-12
claude-opus-4-8
Anthropic
$0.618
$0.618
$11.24
2026-09-12
claude-opus-5
Anthropic
$0.618
$0.618
$11.24
2026-09-12
claude-fable-5
Anthropic
$1.236
$1.236
$22.48
2026-09-18
claude-fable-5-1
Anthropic
$1.236
$1.236
$22.48
2026-09-18
gpt-6-astra
OpenAI
$1.236
$1.236
$22.48
2026-09-17
Read this against the numbers above
Every row below prices the FAST TIER's own 55-call run (r001-coa-conformance) -- the deliberating tier's own token counts are on Cost.cost_by_model[1] and are not separately projected here.
Neither tier's registered run left anything to disable -- src/adapters/__init__.py's thinking parameter is only sent when a caller passes one, and this kit's own harness never does (see LLM.settings) -- so unlike several sibling kits, there is no reasoning-on/reasoning-off discrepancy to caveat here.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Five modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Four of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/segment.pysegment
cut the certificate into addressable sections, pure code
src/segment.py
# Cut a document into addressable sections. Pure code — no model, no network.
def sections(text):
def locate(text, value):
def span_label(secs, start):
src/select.pyselect — a swap seam
pick which sections carry each field, pure code — conformance is mapped to the measured result and the two limits, never to the disposition note
You change it to: map fields to your own certificate's headings; unmatched falls back to the whole document
src/select.py
# Pick which sections plausibly carry each field. Pure code -- the last deterministic step
SECTION_HINTS = {
def for_field(secs, field):
def plan(secs, fields):
src/prompt.pyprompt
assemble one call for all ten fields, with the comparison rule and its boundary-inclusive convention stated in full
src/prompt.py
# Assemble the extraction prompt. One prompt per certificate of analysis, all ten fields in it.
SYSTEM = (
def field_schema(fields):
def build(doc_text, secs, fields, selector):
def parse(raw, fields):
src/extract.pyextract — a swap seam
the AI layer, one provider one key — plus the pure-code self-consistency check downstream: re-run the comparison over the model's OWN extracted numbers and flag any disagreement with the model's own verdict
You change it to: the comparison itself — the boundary convention, or a tolerance band if your programme allows one. It is stated once and used by the corpus generator, the prompt and the guardrail, so changing it here changes all three together
src/extract.py
# Extract one certificate's fields: segment, select, prompt, one model call, then a pure-code
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FIELDS = os.path.join(HERE, "data", "fields.json")
CORPUS = os.path.join(HERE, "data", "corpus")
MAX_TOKENS = 4000
def load_fields():
def load_doc(stmt_id):
def documents():
def recompute_conformance(measured, lower, upper):
def compute(values):
evals/judge.pyjudge
score field accuracy and the conformance confusion matrix separately, pure code
evals/judge.py
# Score an extraction run. PURE CODE -- gold is exact and the answer is one value per cell, so
def norm(v):
def _num(v):
def equal(field, got, want):
def score(fields, records, golds):
def score_flags(records, flags, golds):
Start hereThe shortest path into it
src/segment.pycut the certificate into addressable sections, pure code
src/select.pypick which sections carry each field, pure code — conformance is mapped to the measured result and the two limits, never to the disposition note A swap seam.
src/prompt.pyassemble one call for all ten fields, with the comparison rule and its boundary-inclusive convention stated in full
src/extract.pythe AI layer, one provider one key — plus the pure-code self-consistency check downstream: re-run the comparison over the model's OWN extracted numbers and flag any disagreement with the model's own verdict A swap seam.
evals/judge.pyscore field accuracy and the conformance confusion matrix separately, pure code
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 963 input and 256 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's certificates are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any of the text. In a real deployment the analyst's disposition note is free text written by a person working the batch, and the whole certificate may arrive from a supplier rather than from your own laboratory -- exactly the kind of externally-authored input this kit's architecture reads verbatim and trusts, with no verification step before its text reaches the model. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it — and three of four boundaries hold
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every certificate is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's disposition notes are written by people, and an incoming-goods review reads certificates issued by suppliers, and neither surface has been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it. Confirmed by reading the code, not by a run, on 2026-08-21 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does an extraction ever release, reject or disposition a batch?
An extraction could plausibly post a release decision, a hold or a rejection.
No code path does. src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/extract handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Is the conformance comparison something a prompt or a reply can move?
A crafted disposition note could plausibly talk the arithmetic into a different answer, or shift the boundary convention.
No. recompute_conformance() in src/extract.py is pure code with no configuration and no model in it -- nothing in the reply text is consulted except the four values it names, and needs_review is computed after the model call.
Could a crafted disposition note produce a wrong AND self-consistent answer?
A note written to talk a model into misreading the measured value -- and then judging that misreading correctly -- would satisfy the self-consistency check while being wrong about the batch.
Unmeasured -- no attack has been tried. This is the guardrail's known blind spot (see guardrails.is_not): it compares the reply against itself and never re-reads the certificate. The spannable fields DO carry spans (src/segment.py::locate found 463 of 464 on the fast tier), so the evidence to catch it exists in the record; nothing consumes it. See could_not_verify.
The first three boundaries hold, confirmed by reading the code, not by an attack trial. The fourth is the one this run has not tested, and it is the same blind spot the guardrail names about itself: a consistently-wrong reply looks exactly like a correct one to a check that only compares the reply against itself.
The result0 attack trials, and three of four boundaries checked here hold in code: no write path exists, the comparison is pure code the model cannot move, and a misconfigured key cannot leak into a UI error. The fourth -- whether a crafted note could produce a wrong answer that still agrees with itself -- is unmeasured, and is the guardrail's own known blind spot.
1externally-authored field a live deployment would carry (the analyst's disposition note itself, or a whole supplier-issued certificate) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no certificate's text was authored by an outside party. A real deployment's disposition notes are free text written by whoever worked the batch, and an incoming-goods review reads certificates issued by a supplier entirely outside your control -- exactly the kind of externally-supplied text this kit's architecture reads verbatim and trusts. Whether a note crafted to move conforms_to_spec, or to move the extracted numbers so that a wrong verdict still looks self-consistent, would succeed is unmeasured for this kit.
Read this twice
The guardrail checks the model against itself, not against the certificate. src/extract.py::compute() re-runs the comparison over the numbers the MODEL returned -- so a reply that misreads the measured value and then judges that misreading correctly is self-consistent, and the check stays quiet. That blind spot is the price of the property that makes this guardrail worth having: it needs no gold, so it computes on unlabelled certificates, which is every certificate a forker actually has. On this run it fired 0 times across 110 replies while both tiers scored 55 of 55 verdicts, and it fired 24 times out of 24 errors when run over the free analyst-tone floor's output -- 0 false alarms in both directions. A future version should re-read the three numbers out of the certificate text by pure code before trusting either half, and a red-team pass targeting exactly that blind spot is the natural next measurement.
HonestyWhat this does not prove
Whether a disposition note crafted to talk a model out of its own arithmetic could move conforms_to_spec -- no red-team run exists for this kit. On this corpus's plain register mismatch both tiers were unmoved on all 24 cases.
Whether a crafted note could move the EXTRACTED NUMBERS so that a wrong verdict still passes the self-consistency check -- the check compares the reply against itself and never re-reads the certificate, so by construction it would not notice. Unmeasured.
Whether the live app's own /api/extract behaves identically to the registered run under an adversarial certificate -- both use the same src/extract.py::extract(), but neither has been tested against one.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
needs_review fires when the comparison re-run in pure code over the model's OWN extracted measured_value, spec_lower_limit and spec_upper_limit lands on a different verdict from the model's own conforms_to_spec. It is a SELF-CONSISTENCY check, not a business condition: it needs no gold, no second call and no labelled set, so it computes identically on a certificate nobody has ever scored.
src/extract.py::compute(), which calls recompute_conformance() -- the same comparison tools/build_corpus.py used to write gold and the same one src/prompt.py asks the model to perform, stated once so the three cannot drift. Pure code, run after the model call on the reply's own values; nothing in the reply text can talk it out of firing. What it CANNOT do is check whether those numbers were read correctly in the first place -- a DIFFERENT dependency, named in is_not and could_not_verify.
EvidenceDoes it hold?
What
Measured
The check fires whenever a verdict disagrees with the numbers it was supposedly computed from
Run over the free analyst-tone floor's own output (results/eval-b000-rules.json), it fired on 24 of 55 certificates -- exactly the 24 the floor got wrong, with 0 false alarms and no gold consulted.
The check stays quiet when a reply is internally consistent
0 of 110 replies across both model tiers fired it, and both tiers scored 55 of 55 conformance verdicts against gold -- so on this run quiet and correct coincided.
The comparison is not something a prompt or a reply can move
recompute_conformance() in src/extract.py is pure code with no configuration and no model in it -- confirmed by reading compute(); nothing in the reply text is consulted except the four values it names.
An unanswerable comparison is not a pass
compute() returns needs_review=None, never False, when conforms_to_spec is missing or the measured value is not a number -- confirmed by reading the function, and exercised by the stub run (t000-stub), where all 55 records returned None rather than a clean bill.
No code path releases, rejects or dispositions a batch
src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE NUMBERS ARE RIGHT. It compares the model's verdict against the model's own extracted values; if the model misreads the measured value and then judges that misreading correctly, the reply is self-consistent and this check stays quiet. Catching that needs a second reading of the document, which nothing here does -- see add_first.
It does not decide conformance either way -- the flag says the reply disagrees with itself, not which half is wrong. See UI.shots and Data.breaks_on.
It has not been attacked. Whether a disposition note crafted to talk a model out of its own arithmetic -- or a prompt-injection attempt inside the note field -- could produce a consistent-looking wrong answer is unmeasured; see Security.could_not_verify.
WatchedWhat is watched, and why that one
2runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 15 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
7 measured by the latest run8 need the model half
Metric
Owner
Role
Why this one
field-exact-match-with-normalisation
Per-field exact match, light normalisation
alarm
extraction_accuracy specifically on conforms_to_spec, since that is the field this corpus is built to test — see the confusion-matrix grader below; the two spec-limit fields on the 31 one-sided records, since a model that invents a missing bound scores wrong there AND may then compute a different verdict from it; span_rate on the nine spannable fields, since a value with no span is an assertion rather than a located citation — alarm on Any drop in extraction_accuracy below the measured figures on either tier — this run's near-clean result is the baseline, and any regression means the prompt, the corpus, or the provider changed.
conformance-confusion-matrix
Conformance confusion matrix
alarm
false_negative count specifically — an out-of-specification batch called conforming is the error a quality team cares about most; the self_consistency block alongside it: how often needs_review fired, and how many verdict errors it caught without any gold — that is the only one of these figures a forker can still compute on unlabelled certificates; whether accuracy stays at 1.00 as the corpus grows — 26 positive cases is a small sample — alarm on Any false negative — a certificate gold says is out of specification that the run called conforming.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
55
different corpus — nothing is comparable
corpus.bytes
27,669
certificates edited — the count held, the bytes did not
split.count
495
the sections count moved — a different set was scored
split.size_p50
44
the median size of one section moved
split.size_p95
135
the 95th-percentile size of one section moved
dataset.rows
55
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.002
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (documents 55, extraction_cells 550, failures 0, refusal_cells 0, thinking True) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Extraction accuracy
exact match at 99.82 pct on the fast tier (one product-name misread, COA-0038), 100 pct on the deliberating tier
550 cells
two independent tiers (r001-coa-conformance, r002-coa-conformance) -- one near-exact, one exact -- see Eval.taxonomy.
Conformance accuracy, recall and precision
exact match at 100 pct / 100 pct / 100 pct on both tiers
55 certificates: 26 out of specification, 29 conforming
two independent tiers reproduced the identical figures to the digit.
Self-consistency firings
exact match at 0 on both tiers; 24 of 24 on the free analyst-tone floor
55 replies per tier, 55 baseline records
measured directly on both tiers and on results/eval-b000-rules.json.
Hallucinations
exact match at 0 on both tiers
550 cells
two independent tiers reproduced zero to the digit.
Span rate
99.78 pct on the fast tier (463 of 464 spannable values located), 100 pct on the deliberating tier
spannable extracted values (non-enum fields with a non-null answer; a correctly-returned null on a one-sided specification has nothing to locate)
two independent tiers -- the one gap on the fast tier is the same COA-0038 product-name misread, which locate() could not find verbatim.
Latency
2,415ms / 3,225ms p50/p95 on the fast tier, 4,396ms / 5,336ms p50/p95 on the deliberating tier -- roughly 80 pct higher, not same-input drift, because the two tiers are different models
55 calls per tier
measured directly on both tiers.
Token totals
52,969 input tokens on both tiers (identical prompt); 14,132 output on the fast tier vs 16,243 on the deliberating tier, about 15 pct more
55 calls per tier
measured directly on both tiers.
HistoryRun history
2 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-coa-conformance 2026-08-21
r002-coa-conformance 2026-08-21
extraction accuracy
0.9982
1.0000
invented values
0
0
values with a span
0.9978
1.0000
input tokens, whole run
52969
52969
model latency p50 ms
2415.00
4396.00
model latency p95 ms
3225.00
5336.00
output tokens, whole run
14132
16243
not a time series No two of these 2 runs measured the same system — they differ on the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 2 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
which tier is called (fast vs deliberating)
output tokens (+15 pct) and latency (roughly +80 pct) move together; conformance accuracy, recall, precision, self-consistency firings and hallucinations do not move at all; extraction accuracy and span rate move only very slightly, both on the same single product-name misread
measured
r001-coa-conformance vs r002-coa-conformance: 549/550 vs 550/550 cells, output tokens 14,132 -> 16,243, p50 latency 2,415ms -> 4,396ms, 55/55 verdicts on both.
deciding conformance from the disposition note instead of the numbers
conformance accuracy collapses from 100 pct to 56.4 pct while every structured field stays perfect -- and the self-consistency check fires on every one of the resulting errors, because the numbers it recomputes from are still right
measured
results/eval-b000-rules.json: 495/495 structured cells, 24 of 55 verdicts wrong (10 false negatives, 14 false positives), needs_review fired 24 times with 0 false alarms.
a certificate with a limit on only one side
the comparison drops the missing side entirely -- a model that returns null there computes the same verdict the kit does, and a model that invents a bound may not
reasoning
recompute_conformance() treats a None bound as no constraint -- confirmed by reading the function; 31 of this corpus's 55 records are one-sided and both tiers returned null on every one of them.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Nothing fires yet. Not one of the 7 bands on this kit names a consequence, so there is no row to print. Which is the honest answer to the heading: a band that nothing is wired to is a number somebody has to notice, and this kit has only those. An empty table would have said the same thing while looking like a rendering failure.
NextThe three you would add first
Re-read the three numbers out of the certificate text by pure code and compare them to the model's own extracted values before trusting eitherthe check's blind spot is a consistently-wrong reading (see is_not) -- src/segment.py already does word-boundary matching for spannable fields, and this corpus's numbers are all stated verbatim, so a regex second opinion would close the gap the self-consistency check cannot see.
Compare a batch against its own re-test and its neighbours on the same production runthe verdict is computed from one certificate alone today (see Architecture.breaks_at_scale); a real quality review reads a batch in the context of its own history, and nothing here does that.
Add unit normalisation between the result and the limitevery record in this corpus states both in the same unit by construction, so the kit has never met a ppm result against a percent limit -- Data.breaks_on and could_not_verify both name this as unmeasured.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-check the flag's evidence on any change to recompute_conformance() or compute() in src/extract.py, or to the conforms_to_spec rule in src/prompt.py -- all three change what the check compares. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
What this cannot tell you
Whether the check would still fire cleanly on a model whose extraction errors and verdict errors correlate -- on this run neither tier made a verdict error at all, so the only measured firing evidence comes from the free analyst-tone floor, whose numbers are always right and whose verdicts are often wrong. That is the easy case for this check.
Whether a crafted disposition note could produce a wrong AND self-consistent answer, which this check would not catch -- no red-team run exists for this kit, see Security.could_not_verify.
Whether the specification limits in this corpus resemble any real approved specification -- they do not, by design; see README/SOURCES.md.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is empty, with a comment explaining that the emptiness is load-bearing. The whole extraction decision is three files: src/segment.py, src/select.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
55 certificates of analysis, generated from a fixed seed, never fetched. Gold's conformance label is the arithmetic run over the same numbers the document states, never a typed opinion -- evals/check_labels.py re-derives every one of them before any run is allowed to spend; see data/SOURCES.md.
segmentation and selection
src/segment.py, src/select.py
text splitters / retrievers
a heading-based cut (src/segment.py) and a fixed field-to-heading map (SECTION_HINTS in src/select.py) -- no embeddings, no index, no ranking; a field absent from the map gets the whole document rather than nothing.
prompt assembly
src/prompt.py
prompt templates
the ten-field schema and the conformance comparison, with its boundary-inclusive convention, are one declaration and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. This kit's own MAX_TOKENS=4000 (carried over from the sibling extraction kits' own experience on a similarly-shaped record) is a plain constant, not a client-library setting.
evaluation
evals/judge.py
eval harnesses
per-field exact match plus a separately-scored conformance confusion matrix and a gold-free self-consistency diagnostic is a loop and a handful of counters, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per certificate -- segment, select, prompt, call, parse, recompute -- with no branching and no state carried between certificates. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A different field schema (data/fields.json) or a different comparison rule needs its own gold set and its own evals/check_labels.py pass -- a framework's own schema layer would not remove that work, only relocate it.
SECTION_HINTS is a five-minute edit for a new certificate layout because it is a plain dict, not a configured retriever -- a framework's own chunking/retrieval abstraction would need its own re-tuning pass instead, with its own failure modes to learn.
Swapping providers is one function and one entry in PROVIDERS (src/adapters/__init__.py) -- a framework's own model abstraction would add a dependency and a version to track for the same one-line change this file already gives away free.
What we could NOT verify
Whether a framework's own retrieval or agent abstraction would resolve anything this kit does not already handle was not tested -- this run found one field-level error across both tiers and zero verdict-level errors (see Business.not_good_enough), so there is very little observed weakness here for a framework to address.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-coa-conformance on the fast tier, 2026-08-21. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
2,415 ms
2,415ms / 3,225ms p50/p95 on the fast tier, 4,396ms / 5,336ms p50/p95 on the deliberating tier -- roughly 80 pct higher, not same-input drift, because the two tiers are different models
—
Model, p95
3,225 ms
2,415ms / 3,225ms p50/p95 on the fast tier, 4,396ms / 5,336ms p50/p95 on the deliberating tier -- roughly 80 pct higher, not same-input drift, because the two tiers are different models
—
Input tokens
52,969
52,969 input tokens on both tiers (identical prompt); 14,132 output on the fast tier vs 16,243 on the deliberating tier, about 15 pct more
—
Output tokens
14,132
52,969 input tokens on both tiers (identical prompt); 14,132 output on the fast tier vs 16,243 on the deliberating tier, about 15 pct more
—
No movement column. Not one of the 1 earlier run on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-coa-conformance2,415 ms
r002-coa-conformance4,396 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-21, across 2 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
certificates of analysis
data/corpus/*.txt — 55 files, generated once from a fixed seed (SEED = 20260821) by tools/build_corpus.py
read whole by src/segment.py and src/select.py; never modified after generation
gold labels
data/gold.jsonl — 55 rows, whose conforms_to_spec is the arithmetic run over the same numbers the document states, never a typed opinion, checked against the document text by tools/build_corpus.py's own _verify() pass and re-derived by evals/check_labels.py before any run may spend
never — evals/judge.py is pure code, no model, no key
the field schema
data/fields.json — the ten-field record src/prompt.py assembles the user message from
read by src/prompt.py and src/extract.py only
the key
.env — never committed (see .gitignore)
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
model
one call per CERTIFICATE, carrying the selected sections plus the ten-field schema, behind src/adapters/__init__.py — OpenAI-compatible wire format over raw HTTP, so a forker runs this on whichever key they hold. MAX_TOKENS is fixed at 4000 (src/extract.py), set on the sibling extraction kits' own experience with a similarly-shaped JSON record rather than a ceiling this kit has ever hit — 0 of 55 calls truncated on either tier.
2,415ms p50 / 3,225ms p95 on the fast tier, 4,396ms p50 / 5,336ms p95 on the deliberating tier (the fast-tier and deliberating-tier runs, 2026-08-21 -- see Cost.cost_by_model)
one call per certificate, no concurrency and nothing shared between calls — see Architecture.breaks_at_scale. A release queue of thousands needs batching and a rate-limit strategy this kit does not have; MAX_CALLS_PER_DAY in src/budget.py caps the shared key across every kit on this machine, not this kit's own throughput.
point src/adapters/__init__.py at a different provider or model and every published accuracy/cost/latency figure is void until re-run — this run's numbers are this model's numbers, not a property of the prompt.
corpus refresh
tools/build_corpus.py regenerates the whole corpus — 55 certificates across a fixed roster of invented products, six test parameters and eight disposition notes — byte-identically from a fixed seed (SEED = 20260821) every time it is run. There is no incremental refresh; gold's conformance label is the arithmetic run over the same numbers the document states, and the script's own _verify() pass checks that every stated value appears verbatim in the document text and that every label agrees with its own numbers before anything is scored.
0.038s wall time to regenerate all 55 files and their gold labels (measured directly, 2026-08-21 (python3 tools/build_corpus.py, timed) -- see Data.index for the separate segmentation figure, a different step)
a real plant's batch volume, product roster and certificate layouts do not come from a fixed seed and grow without bound — how a genuinely varied set of real certificates would change what src/segment.py (heading-based) and src/select.py (SECTION_HINTS) can resolve was not measured; see Architecture.breaks_at_scale and Data.breaks_on.
point tools/build_corpus.py at your own certificates and specification library, and every published accuracy figure is void — they are this corpus's own planted ambiguity, not a property of the model.
labels
evals/check_labels.py asserts gold/document consistency (one gold row per document and vice versa, every enum value allowed by the schema, no record with a limit missing on both sides) AND re-derives every conformance label from its own numbers before evals/run.py is allowed to spend anything. Scoring (evals/judge.py) is per-(certificate,field) exact match with light normalisation, plus a separately-scored conformance confusion matrix and a self-consistency diagnostic computed from the model's own extracted fields by pure code (src/extract.py::compute), never from a second model call.
550 of 550 possible (certificate, field) cells scored (55 certificates x 10 fields) plus 55 certificates carrying a conformance verdict and a self-consistency verdict, scored separately (lenses.Eval.dataset, the fast-tier and deliberating-tier runs, 2026-08-21)
the gold set stops at 55 certificates and covers one planted ambiguity (a disposition note whose tone contradicts the arithmetic) at a 44 pct achieved rate — a real quality queue's mix of confusable phrasing is unmeasured, and a unit mismatch or a superseding re-test is never presented; see Data.breaks_on.
a different field schema (data/fields.json) or a different comparison rule needs its own gold set and its own evals/check_labels.py pass before any published figure can be trusted again.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
conforms_to_spec answered 'no' on a certificate whose analyst note says the batch was within normal range and released
the model did the comparison rather than being talked out of it by the reassuring prose in the same document — this is the corpus's planted ambiguity (AMBIGUOUS_FRACTION=0.40 in tools/build_corpus.py, 44 pct achieved on this seed), and on this run both tiers got every one of these cases right (see Eval.taxonomy)
read the three numbers before assuming a 'no' next to a reassuring note is a mistake — the prompt's own rule (see README) is that the arithmetic decides and the note is an opinion, and this run found zero cases where either tier took the prose shortcut instead (the fast-tier and deliberating-tier result files, both runs, 2026-08-21 -- see Eval.taxonomy and Eval.baseline for the same pattern's effect on the free analyst-tone floor, which took the shortcut on all 24)
needs_review fired on a certificate
the reply disagrees with itself: the comparison re-run over the model's OWN extracted measured value and limits landed on a different verdict from the one the model gave. One of the two halves is wrong and neither is trustworthy until somebody looks. It fired 0 times across 110 replies on this run, and 24 times out of 24 errors when run over the free analyst-tone floor's output
compare the three numbers on the panel against the certificate itself — the check knows the reply is inconsistent, not which half is wrong (src/extract.py::compute over both runs' result files, and over results/eval-b000-rules.json, 2026-08-21)
No machine symptom — this failure leaves no trace in any output.
no path in src/extract.py::compute() or src/app.py releases, rejects or dispositions a batch — conforms_to_spec and needs_review are returned to the caller as fields on the response record and nothing downstream of this kit acts on them. A wrong verdict on a real deployment would show up only in whatever quality system consumes this kit's output, which this kit does not have and does not simulate — so there is no committed artifact naming that failure, on purpose: it is out of this kit's boundary, not unmeasured.
Whether a 55-certificate, single-seed run's clean result generalises to a harder register mismatch this corpus did not plant, or to a larger sample, was not tested — see Eval.could_not_verify. Concurrency (every run in this series is one call at a time, sequential), unit conversion between a result and a limit stated differently, a specification that changed between test and review, and how either tier performs on a real multi-parameter certificate with a results table are all unmeasured — see Eval.could_not_verify and Data.breaks_on for the full list.
The corpus licence, from the Data lens: MIT — this repository's own licence. Every product, batch, analyst and specification limit is invented; a real certificate of analysis cannot be published either — it names a real manufacturer, a real lot and a real release decision. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one linePer-field exact match, light normalisation
Does the model's value for each field match gold, after trimming whitespace/punctuation and treating numbers within half a cent as equal — including the two limits, where a correctly-returned null on a one-sided specification is a hit and an invented bound is a miss?
$0.00per 1,000 certificates
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score, in-process, no key and no model. The same function evals/run.py calls after every call, and the same field-match logic the free baseline is scored by — a baseline and a real run scored by two different scorers cannot be compared honestly.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The certificate
COA-0033
The field this row is about
conforms_to_spec
What the laboratory measured
1025
In
CFU/g
Lower specification limit
not stated — this is a one-sided specification
Upper specification limit
1000
What the model answered
no
What the arithmetic says
no
Recomputed from the model's own numbers
no
Self-consistency check fired
no
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
COA-0033 states a total plate count of 1025 CFU/g against a one-sided specification — no lower limit, upper limit 1000 — and an analyst note reading "Within normal range for this product line. Released to finished-goods inventory." Both tiers returned all ten fields exactly, including spec_lower_limit as null rather than inventing a bound.
Conformance confusion matrix
correct
Gold conforms_to_spec=no, derived by comparing 1025 against an absent lower limit and an upper limit of 1000. Both tiers answered no despite the reassuring note, and src/extract.py::compute() re-ran the comparison over each tier's own extracted numbers and agreed — needs_review stayed false on both. The free analyst-tone floor answered yes here, one of its 10 false negatives.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 99.8%
the deliberating tier
scored 100.0%
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold, read back off the same numbers, note and dates the document states; evals/check_labels.py asserts every non-nullable field is populated on every row, that no record states a limit on neither side, and that every conformance label agrees with its own arithmetic, before any run is allowed to spend.
These rates are UNKNOWN, on purpose
This grader IS the reference standard for field values — it cannot be scored against itself. What can go wrong is the corpus's own generation logic, which tools/build_corpus.py's own _verify() pass checks by confirming every stated value appears verbatim in the document and every date parses as a real date.
Watch these
extraction_accuracy specifically on conforms_to_spec, since that is the field this corpus is built to test — see the confusion-matrix grader below
the two spec-limit fields on the 31 one-sided records, since a model that invents a missing bound scores wrong there AND may then compute a different verdict from it
span_rate on the nine spannable fields, since a value with no span is an assertion rather than a located citation
Alarm on
Any drop in extraction_accuracy below the measured figures on either tier — this run's near-clean result is the baseline, and any regression means the prompt, the corpus, or the provider changed.
How tight can the band be? There is no tolerance band on the field grade itself — it is exact match after trimming whitespace/punctuation, and numbers within half a cent are treated as equal, never a continuous score to round.
Cadence: Re-score on any change to src/prompt.py, data/fields.json or tools/build_corpus.py — the first two change what is asked, the third changes what is asked about. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
The decisionWhen to reach for it
Use it
Gold is read back off the same numbers and text the certificate states, never from a separate target — true of every kit corpus, never true of a real quality team's own certificate archive.
Do not use it
The true field values are not known in advance — the normal state of a real release review, and the reason this corpus is generated rather than captured.
Check each production batch's lab certificate against its limits
PresenterOpens the private repo. Visible to admins only.
In one lineConformance confusion matrix
Does the model's stated conforms_to_spec match the same comparison run over gold's own measured value and limits? Out of specification is the positive class, so recall answers "of every batch that really is out of specification, how many did the run call out of specification".
$0.00per 1,000 certificates
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model. Derives gold's verdict with src/extract.py::recompute_conformance, compares it to the model's own conforms_to_spec per certificate, and reports the self-consistency diagnostic (src/extract.py::compute) alongside it.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The certificate
COA-0033
The field this row is about
conforms_to_spec
What the laboratory measured
1025
In
CFU/g
Lower specification limit
not stated — this is a one-sided specification
Upper specification limit
1000
What the model answered
no
What the arithmetic says
no
Recomputed from the model's own numbers
no
Self-consistency check fired
no
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
COA-0033 states a total plate count of 1025 CFU/g against a one-sided specification — no lower limit, upper limit 1000 — and an analyst note reading "Within normal range for this product line. Released to finished-goods inventory." Both tiers returned all ten fields exactly, including spec_lower_limit as null rather than inventing a bound.
Conformance confusion matrix
correct
Gold conforms_to_spec=no, derived by comparing 1025 against an absent lower limit and an upper limit of 1000. Both tiers answered no despite the reassuring note, and src/extract.py::compute() re-ran the comparison over each tier's own extracted numbers and agreed — needs_review stayed false on both. The free analyst-tone floor answered yes here, one of its 10 false negatives.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 100.0%
In operationWhat to monitor
Reference standard: The same pure-code comparison in src/extract.py::recompute_conformance, run over gold's own measured_value, spec_lower_limit and spec_upper_limit rather than the model's — the TRUE verdict is never separately typed, only derived.
These rates are UNKNOWN, on purpose
Whether the verdict holds on confusions this corpus did not build — a unit mismatch, a limit expressed relative to a target, a superseding re-test. The matrix is exact against this run's 55 certificates and says nothing beyond them. See Eval.could_not_verify.
Watch these
false_negative count specifically — an out-of-specification batch called conforming is the error a quality team cares about most
the self_consistency block alongside it: how often needs_review fired, and how many verdict errors it caught without any gold — that is the only one of these figures a forker can still compute on unlabelled certificates
whether accuracy stays at 1.00 as the corpus grows — 26 positive cases is a small sample
Alarm on
Any false negative — a certificate gold says is out of specification that the run called conforming.
How tight can the band be? No tolerance band — a binary verdict either matches gold's derived verdict or it does not, scored per certificate.
Cadence: Re-score on any change to src/extract.py::recompute_conformance or to the conforms_to_spec instruction in src/prompt.py — both change what the verdict means.
The decisionWhen to reach for it
Use it
Gold's true verdict is derived by running the same comparison the kit itself publishes, over gold's own numbers — never a separately-typed truth that could drift from the rule the kit actually applies.
Do not use it
A real release decision is not known until a qualified person signs against an approved specification — this corpus's gold is a construction, not an observed outcome.
A living map of modern AI — kept current every morning