The business caseThe problem this solves
One laboratory result comes back out of specification, and before anything can be decided about it somebody has to assemble the evidence FILE behind it: the original result record, the phase-1 laboratory check, the governing retest, the sister-lot trend comparison, the instrument's calibration and maintenance history inside a declared lookback window across a declared pair of systems, the analyst's qualification and the method version in force on the test date. Eight elements, each either present WITH the record that satisfies it or missing for a stated reason. The records are printed on the packet with STATUS columns, so most of the job is a column read — and on 21 of the 64 packets in this corpus a NOTE somewhere else in the file takes a printed record out of play, or looks exactly as if it does and does not. Opening one out-of-spec evidence packet, reading the five printed record blocks against the eight-element checklist, testing each candidate record against the packet's own DECLARED lookback window and declared pair of instrument systems, reading every note to see whether it takes a record out of play AS AT THE RESULT DATE or merely mentions it, and carrying the sister lots, instruments and analysts the cited records drag in out with the answer.
Audience
The quality-systems desk of a manufacturer assembling the evidence file behind one out-of-spec result, and the laboratory, metrology and training desks whose records it cites. The decision it supports is 'is this file complete enough to hand on' — never what caused the result, never what happens to the batch. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual evidence packets
The corpus is 64 evidence packets, 0.10 MB (json 3 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real out-of-spec investigation file is a manufacturer's own regulated record: named analysts on every result, product and batch identity, instrument serials, and laboratory-desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that takes a record out of play — stripped out of it first. So the whole thing is invented, declared synthetic in data/SOURCES.md, and generated from SEED 20260914 with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 evidence packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every product, batch/lot, instrument, analyst code, method and record id is invented, generated from SEED 20260914, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL - an analyst is a code and every note speaks for the laboratory, quality, metrology, training, documentation, production, planning or supply desk. evals/check_labels.py sweeps all 64 packets for a person-shaped name, an honorific and an unattributed note on every run and reports 0. The best free-code floor is published there too: 43 of 64 whole-packet, measured before any call was bought.
Swap this folder for your own material and the kit is pointed at your evidence packets. That is the whole change — there is no database to migrate.
OUT-OF-SPEC EVIDENCE PACKET OOS-0001
PACKET ASSEMBLED 2026-05-16 STANDARD OOSE-2026
PRODUCT PRD-4407 STRENGTH 750 MG
BATCH/LOT LOT-26-0400
TEST ELEMENTAL IMPURITIES METHOD MTH-148
RESULT DATE 2026-05-06 INSTRUMENT ICPM-5529 ANALYST ANL-3144
SPECIFICATION 95.00 - 105.00 % REPORTED 92.67 % OUT OF SPEC
DECLARED SCOPE SYSTEMS CALMGR, EAMS LOOKBACK 180 DAYS FROM 2025-11-07
[1] LABORATORY RECORDS
ID DATE TYPE ANALYST STATUS
LIR-70000 2026-05-06 ORIGINAL-RESULT ANL-3144 REPORTED
LIR-70001 2026-05-07 PHASE-1-CHECK ANL-3144 COMPLETE
[2] RETEST RECORDS
ID DATE AUTHORISATION ANALYST VALUE STATUS
RT-30000 2026-05-08 RTA-9699 ANL-3173 97.97 % REPORTED
[3] SISTER-LOT TREND ROWS
ID LOT PRODUCT STRENGTH TESTED VALUE
TRD-61000 LOT-26-0100 PRD-4407 750 MG 2026-03-29 100.91 %
TRD-61600 LOT-26-0752 PRD-4407 250 MG 2026-03-15 102.00 %
[4] INSTRUMENT HISTORY
ID SYSTEM INSTRUMENT TYPE DATE OUTCOME
CAL-88000 CALMGR ICPM-5529 CALIBRATION 2026-04-13 PASS
MNT-55000 CALMGR ICPM-5529 MAINTENANCE 2026-02-11 CLOSED-NO-DEFECT
[5] REFERENCES
ID KIND SUBJECT VALID FROM VALID TO STATUS
QUA-44000 QUALIFICATION ANL-3144 2025-05-05 2027-05-02 SIGNED
MVR-11000 METHOD-VERSION MTH-148 2025-09-30 2027-07-17 REISSUED
[6] NOTES
Laboratory desk: no note was raised on this packet.
The outcomeWhat a good result looks like
One evidence packet in, one row out: the batch/lot, all eight OOSE-2026 element rows with the record ids that satisfy each or the reason it is missing, the missing list, one COMPLETE or INCOMPLETE verdict about the FILE, and the three lineage lists — sister lots compared, instruments touched, analysts involved. 62 of 64 packets come back with all seven graded fields right, against 43 for the floor of record and 37 for a naive vocabulary rule.
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 2 packets it got wrong are named in the kit README with the sentence that decided each, and both are the voided reading: OOS-0012 missed a void that was there, OOS-0007 acted on one dated after the result date. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your laboratory system records a supersession, a void or a suspended qualification as a STATUS on the record itself, and your instrument history is exported for the declared systems only — the free columns floor, and do not buy a call at all
43 of 64 packets whole for $0.00, and on the 43 packets the columns decide it is 43 of 43 while the paid arm is 42. Scope, window, instrument match and the two-sister-lot rule are all decidable from printed rows and dates. - Record changes arrive as free text — a laboratory-desk note, a training-desk line, an email pasted into the file — and some of them are decoys — the paid call
This is the whole product. 25 of 25 decoy notes read correctly against the best floor's 18, and 20 of 21 notes that really void a record against 15. - Your notes routinely carry dates, and a note dated after the result date must decide nothing — the free columns floor, and read the paid arm beside it
4 date traps: the columns floor gets all four for the trivial reason that it reads no note, the paid arm gets three. This is the one family where the paid call is measured WORSE and it is named on the kit's own front page. - You need the sister lots, instruments and analysts a file drags in — either — and the honest answer is the free one
64 against 59, five discordant packets in one direction, p = 0.0625 against BOTH free floors. Not statistically distinguishable on this corpus.
And where nothing here is good enough:
- Untrusted text can reach the notes block — a supplier portal, an email ingest, an OCR of a scanned annotation — neither, until you have closed that path
One planted sentence moved the void reading on 7 of 7 packets. The cap held 21 of 21 and is enforced as a SHAPE, but the reading is not defended by anything in this kit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own evidence packets in the same shape — header, DECLARED SCOPE line, five printed record blocks, notes — and data/packets.json with your own register, then edit data/policy.json so the eight elements, the nine reason codes and the per-element ladders are YOURS. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per packet for arithmetic you already have. That is the case against the best-fitting scenario (“Your laboratory system records a supersession, a void or a suspended qualification as a STATUS on the record itself, and your instrument history is exported for the declared systems only”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet with no DECLARED SCOPE line. The whole instrument-history reading is 'inside the declared systems and the declared window'; with neither declared there is nothing to test against and the kit would be inventing a scope, which is the one thing the candidate row's open question says not to do. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-oos-evidence. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, runs all three free floors live in the browser and replays every committed arm. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run; without one the button is disabled and says so.








