Check who's liable when a retailer's delivery doesn't add up
Somewhere on a delivery, one line doesn't match across the purchase order, the bill of lading and the dock count. This app finds that line, checks which item any carrier note names, and drafts a notice proposing who is liable for you to confirm.
PresenterOpens the private repo. Visible to admins only.
For receiving and accounts payableCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
Receiving and accounts payable staff at a retailer, building the case file before any chargeback or credit goes out.
✕Today's manual process
1Line up three counts for every line: what was ordered, what the bill of lading says, what the dock scanned.
2Find the one line that doesn't agree, among the lines that do.
3Check any carrier note to see which item it actually names, then decide who is liable.
4One wrong call means a chargeback the documents don't support.
Every line checked by a person
✓With the app
1Every line is read with its ordered, shipped and received counts.
2The line that doesn't agree is found, with the two counts that differ and the gap.
3A carrier note counts only for the item it names; when the papers don't settle it, the app says so.
4A liable party is proposed with a drafted notice, and your team confirms before any chargeback or credit.
People confirm a drafted proposal
See it work
One real case: what the app reads, step by step
A delivery carried by Union Point Freight: one line ordered at 105 was shipped and received at 124, and no carrier note is on file.
Check who's liable when a retailer's delivery doesn't add upReference appBuilt to be shaped to your process
6
1The case one delivery: its purchase order, its carrier and its two lines.
2A line that agrees 334 ordered, 334 shipped, 334 received.
3The line that doesn't 105 ordered, but 124 shipped and 124 received.
4No carrier note nothing on file for this delivery.
5Who is liable insufficient evidence, since the papers and any carrier note don't settle it, with why.
6The proof the line that disagrees, its two counts and the gap, filled in when the check runs.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check who's liable when a retailer's delivery doesn't add up
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A receiving discrepancy case can carry 2-4 SKU lines and exactly one is genuinely discrepant -- its PO, BOL and receiving quantities do not all agree. A carrier exception note may or may not be on file, and when one is, it is only evidence for the SKU it actually names, not for the case in general. Reading 'an exception exists somewhere' as 'the carrier did it' produces a liable-party call the documents do not support. Receiving/AP staff reading a receiving-discrepancy case's PO, BOL and warehouse-scan quantities line by line, checking whether any carrier exception note on file actually names the discrepant SKU before proposing who is liable -- for every case, every day.
Audience
Receiving and AP staff who build the case file before a chargeback or credit is proposed, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual cases
The corpus is 36 cases, 0.01 MB (jsonl 1). A real receiving discrepancy names a real vendor, a real carrier and a real chargeback dollar figure -- this repo has never held one and is not able to. There is no public corpus of paired (BOL, receiving scan, carrier exception, adjudicated liable party) records, for the same reason no public corpus of supplier agreements or production alerts exists: the interesting cases are exactly the ones nobody can publish. Nothing in the corpus refers to a real company -- SKUs, PO numbers, case IDs and carrier names are all invented.
The corpus
The 36 casesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your cases. That is the whole change — there is no database to migrate.
One case, as the model receives itcases.jsonl · 1 of 36
A proposed liable party (vendor, carrier, internal, or insufficient_evidence), the discrepant SKU, the two quantities and delta the notice cites, and a drafted notice -- for receiving/AP to confirm before any chargeback or credit is issued.
And when it cannot
Crediting a carrier exception note without checking which SKU it actually names (wrong_sku_exception_misread), or forcing a confident vendor/carrier/internal call on a case the evidence does not cleanly support (false_confident_call) instead of saying so. Measured at 0 of 6 planted trap cases and 0 of 6 gold insufficient_evidence cases this run -- see not_good_enough for why zero on one run is not the same claim as zero on a larger or adversarial sample.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Resolving a receiving-discrepancy case's liable party from its PO/BOL/receiving quantities and any carrier exception note — the fast tier, over the free exception-first floor 100% liable_party accuracy and 0% wrong_sku_exception_misread this run, against the floor's 83.3% and 100% wrong_sku_exception_misread rate on the exact cases built to be misread on exception presence alone.
Deciding whether a case has a clean documentary basis for a liability call at all — the fast tier's liable_party verdict, read alongside whether it answered insufficient_evidence rather than forcing a guess 0 of 6 gold insufficient_evidence cases this run were given a confident vendor/carrier/internal call (false_confident_call=0) -- the expensive-direction error this kit's guardrail exists to catch.
What this kit is not — a proposed liable party and its documentary basis for receiving/AP to review -- never a chargeback, never a credit, never a final determination src/resolve.py and src/app.py have no function anywhere that issues a chargeback, posts a credit, or records a final liability decision.
At a glanceHow the whole thing runs
100%liable accuracy pct
4,090 msp50, end to end
$2.14per 1,000 cases · Google Gemini 3 Flash
Run once, for real, on 2026-08-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check who's liable when a retailer's delivery doesn't add up14 steps · 4 questions · run once, for real · 2026-08-20
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py's CASE_PLAN, CARRIERS and EXCEPTION_TEMPLATES at your own SKUs, carriers and case shapes, or write your own data/cases.jsonl in the same shape -- src/resolve.py and src/prompt.py read a case by its own fields (lines[].sku/po_qty/bol_qty/received_qty, carrier_exception) and do not care where the numbers came from. The measured 100% liable_party accuracy and 0% false_confident_call/wrong_sku_exception_misread rates are THIS corpus's exception-note phrasing (6 hand-written templates) and THIS corpus's trap design (6 of 14 internal cases, ~43%, planted on purpose).Corpus lens →
When is this the wrong choice?
Avoid: The exception-first floor for anything beyond the non-trap cases, where blaming the carrier on any exception happens to agree with the correct call -- it is not a competitor, it is the honest floor a model has to clear. That is the case against the best-fitting scenario (“Resolving a receiving-discrepancy case's liable party from its PO/BOL/receiving quantities and any carrier exception note”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
This kit assumes every case's documentary evidence (BOL, receiving scan, carrier exception) is complete -- a real deployment where a document was never captured (a BOL nobody scanned, a receiving count nobody logged) has no way to be told apart from one that never existed, and this kit has no way to notice the trail it is reading is incomplete rather than merely unclear. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether this run's 100% liable_party accuracy and 0% false_confident_call/wrong_sku_exception_misread are stable properties of this model at these settings, or a one-off -- one recorded run is not a history. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-20 — r001-rcv-disc. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- every case's lines and its carrier exception note, since data/cases.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Clicking Resolve with no API_KEY returns a calm 200 explaining nothing was called (src/app.py's /api/check). It cannot reproduce a determination, a score, or a dollar figure without a key -- those are what results/eval-r001-rcv-disc.json already committed.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
4,090 msp50, end to end
9,869 msp95
2 minclone to first result
What the clock covers. END-TO-END per case: one HTTP request carrying every SKU line (PO/BOL/receiving quantities) and the carrier-exception note (or its absence), parsed to a liable party, discrepant SKU, cited quantity pair, delta and drafted notice. No retrieval step -- loading a case from disk is pure code and costs no network time. This run left provider-side reasoning ('thinking') at its documented default (on) rather than disabling it, so p50/p95 include a hidden reasoning pass on every call, not only the JSON-formatting cost -- see not_good_enough.
Current processWhat it replaces
Receiving/AP staff reading a receiving-discrepancy case's PO, BOL and warehouse-scan quantities line by line, checking whether any carrier exception note on file actually names the discrepant SKU before proposing who is liable -- for every case, every day.
Where it is not good enough
Three real findings. First, a live discrepancy between what this report measures and what a forker's own click actually does: the registered run (r001-rcv-disc) left provider-side reasoning ('thinking') at its documented default -- ON -- while the shipped app's own endpoint (src/app.py's /api/check) hardcodes thinking=THINKING_OFF on every live call. Nobody reconciled the two before this run was registered as the measured fact. Reasoning consumed 76.9% of this run's total output-token budget (15161 of 19706 output tokens; 421.1 tokens/call average, up to 1206 on one case) and is a real driver of both this kit's latency (p50 4090ms, p95 9869ms) and, since a provider bills reasoning tokens as completion tokens, its dollar cost -- the published cost and latency numbers on this report are therefore NOT what a reader gets from the shipped UI, that configuration has not been separately measured. Second, this run's own 100% liable_party accuracy and 0% false_confident_call/wrong_sku_exception_misread rates are measured on one 36-case run, not a distribution -- no repeat exists (see Eval.could_not_verify), and wrong_sku_exception_misread's own denominator is exactly 6 planted trap cases, which is what a corpus of this scale honestly supports (see data/SOURCES.md) rather than a stable rate over a larger sample. Third, a structural limitation stated in the corpus's own documentation, not found by this run: this kit assumes every case's documentary evidence (BOL, receiving scan, carrier exception) is complete -- a document a real deployment failed to capture (a BOL nobody scanned, a receiving count nobody logged) looks, to this kit, identical to a document that never existed, and this kit has no way to notice the trail it is reading is incomplete rather than merely unclear; every case in this corpus has complete evidence by construction.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free exception-first baseline (evals/baseline.py, blames the carrier on any exception note without checking which SKU it names) scores 83.3 pct liable-party accuracy overall but misclassifies all 6 planted trap cases — a 100 pct wrong_sku_exception_misread rate — exactly the wrong-SKU trap this kit's model resolved correctly on all 6 planted instances. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
whether reasoning is enabled
src/resolve.py and src/app.py
check() takes a thinking kwarg; src/app.py's live /api/check hardcodes it OFF while the registered eval run (r001-rcv-disc) left it at provider default (on) -- a real, live discrepancy between what this report measures and what a forker's own click experiences today. See Business.not_good_enough.
A fifth liable-party value, or a different precedence rule, needs a rewrite of LIABLE_PARTIES/LIABLE_MEANINGS plus classify_case()'s own logic -- a prompt tweak alone will not teach a rule the corpus never plants.
the corpus
tools/build_corpus.py
Point it at your own SKUs, carriers and case shapes. src/resolve.py and src/prompt.py read a case by its own fields (lines[].sku/po_qty/bol_qty/received_qty, carrier_exception) and do not care where the numbers came from.
the trap fraction
tools/build_corpus.py (CASE_PLAN)
Set by a six-category plan tuple, not a single global -- a new corpus states its own category counts in the same shape.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 36 receiving-discrepancy cases from a fixed seed (SEED=20260819) -- 2-4 SKU lines per case, one genuinely discrepant, plus an optional carrier exception note naming exactly one SKU. Gold (liable_party, discrepant_sku, doc_a/qty_a, doc_b/qty_b, delta, trap) is RECOMPUTED from the actual generated case record by classify_case(), never from the target category the generator aimed for -- --verify independently re-checks every row against the record actually written to disk.
the prompt
src/prompt.py
The four-value liability taxonomy and the citation vocabulary, declared once and read from here by build() and parse(). States the wrong-SKU-exception failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover.
the AI layer
src/resolve.py
Loads one case by id, calls the model once with every SKU line and the carrier exception note (or its absence), parses the reply into a liable party, discrepant SKU, cited quantity pair, delta and drafted notice. The whole AI layer, deliberately short -- the same split data-reconcile's reconcile.py and fin-payrun's payrun.py both make. Never issues a chargeback or credit and never makes the final liability call -- there is no function here or in src/app.py that does either.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
spend control
src/budget.py
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
the app
src/app.py
Static files and three JSON endpoints (/api/state, /api/case, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-rcv-disc) actually used; see Business.not_good_enough.
the scorer
evals/scoring.py
Four-way liable_party exact match, pure code: false_confident_call (a confident call on a gold insufficient_evidence case) and wrong_sku_exception_misread (a trap case credited to the carrier anyway) reported as their own named figures rather than folded into a generic 'wrong' bucket, plus discrepant_sku and quantity-citation accuracy checked separately.
free baseline
evals/baseline.py
A fixed rule, written the way a person free-texting a quick classifier would write one, that blames the carrier on ANY exception note without checking which SKU it names -- so it misclassifies every planted trap case by construction. The honest, narrow floor a model has to clear.
Where it breaks at scale
Not on case count -- each case is independent and one call per case is linear. It breaks two other ways. First, the STRUCTURAL ASSUMPTION: this kit assumes every case's documentary evidence (BOL, receiving scan, carrier exception) is complete (see Data.breaks_on) -- a document a real deployment failed to capture has no way to be told apart from one that never existed; stated plainly in data/SOURCES.md, not discovered by this run. Second, REASONING LEFT ON WELL UNDER THE CEILING: MAX_TOKENS is fixed at 3000 (src/resolve.py), sized generously up front rather than computed -- this run's worst observed call used 1,322 of the 3,000-token budget (1,206 of it reasoning), well under half, but reasoning consumed 76.9% of the average output-token budget and was left at provider default (on) rather than disabled -- a call with meaningfully more reasoning than this run's worst case was never measured against the ceiling.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Case CASE-00001 -- PO PO-356479, carrier Union Point Freight: both lines shown (SKU-4001 clean at 334/334/334, SKU-4002 discrepant at PO 105 / BOL 124 / Received 124), no carrier exception on file, before any call is made.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same case after pressing Resolve with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. No proposed determination is shown -- this is the honest failure state, not a staged one.failureOpen full size →
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
36cases
0.01 MiBjsonl 1
p50 283chars per characters per case's assembled block
$0.00setup · 0.0s
How it is cutWhat one characters per case's assembled block is
One case, its 2-4 SKU lines and its one carrier-exception field (or none) included whole -- no train/test split, every case resolved once, in one call.
SetupWhat the setup figure measured
No index is built. src/app.py finds a case by scanning R.cases() for a matching case_id -- 36 rows, a linear scan, not a search structure -- and src/resolve.py sends that case's full record whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own cases
Point tools/build_corpus.py's CASE_PLAN, CARRIERS and EXCEPTION_TEMPLATES at your own SKUs, carriers and case shapes, or write your own data/cases.jsonl in the same shape -- src/resolve.py and src/prompt.py read a case by its own fields (lines[].sku/po_qty/bol_qty/received_qty, carrier_exception) and do not care where the numbers came from. The four-value taxonomy in src/prompt.py assumes exactly these four liable parties; a different set is a prompt.py AND tools/build_corpus.py's classify_case() change, not a data-only change.
⚠︎ And what stops being true when you do: The measured 100% liable_party accuracy and 0% false_confident_call/wrong_sku_exception_misread rates are THIS corpus's exception-note phrasing (6 hand-written templates) and THIS corpus's trap design (6 of 14 internal cases, ~43%, planted on purpose). A real carrier's own exception-note phrasing, and a real receiving operation's own missing-document shape (see breaks_on), are both unmeasured by this run.
What breaks it
This kit assumes every case's documentary evidence (BOL, receiving scan, carrier exception) is complete -- a real deployment where a document was never captured (a BOL nobody scanned, a receiving count nobody logged) has no way to be told apart from one that never existed, and this kit has no way to notice the trail it is reading is incomplete rather than merely unclear.
One fixed liability-attribution rule, applied to every case. A real vendor/carrier relationship's actual agreement can vary who bears transit risk and at what point title transfers; this corpus applies the single four-way rule in classify_case() to every case, the same way data-reconcile documents its own per-supplier limitation.
At most one discrepant line per case. A real receiving dock can find more than one line short on the same shipment at once; this corpus plants exactly one per case to keep the failure mode legible.
The trap denominator is small. 6 trap cases (of 14 internal cases) is what a corpus of this scale supports honestly -- a wider real deployment would want a larger sample before trusting the wrong_sku_exception_misread rate as a stable number.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,117
915
case
195
57
Total
972
This is the cost lesson as arithmetic: of the 972 tokens assembled, 915 are instructions — 94% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged verbatim by importing src.prompt and src.resolve directly and calling prompt.build(case) for CASE-00001 (not retyped) -- the literal system and user content src/resolve.py::check() sends. Per-part token counts are a proportional estimate over character share (system 3117 / case 195 chars, 3312 total) applied to this call's own total of 972 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split. The 'case' part covers only the case's lines and exception-note block src/prompt.py::_case_block() builds -- the surrounding 'Build the case file...' instruction text is part of the user message actually sent (see prompt_verbatim) but is not a separately named segment in the kit's own code.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You build a receiving-discrepancy case file for one case and propose a probable liable party, for receiving/AP to confirm before any chargeback or credit is issued. You are given every SKU line on the case -- each with the quantity the PO ordered, the quantity the BOL states was shipped, and the quantity receiving actually scanned -- plus one carrier exception note, which is either absent or names ONE specific SKU.
Exactly one line on the case is genuinely discrepant: its PO, BOL and receiving quantities do not all agree. Every other line is clean (all three quantities equal) -- find the discrepant line among them, do not assume it is the first one listed.
THE FOUR VALUES FOR liable_party, EXACTLY ONE APPLIES
vendor the BOL states LESS than the PO ordered, and receiving recorded exactly what the BOL says was shipped. The carrier delivered everything it was given; the vendor gave it less than the PO called for. Cite PO vs BOL.
carrier the BOL states the full PO quantity was shipped, receiving recorded less than that, AND the carrier's own exception note names this exact SKU. The carrier picked up the full order and delivered less, and its own paperwork corroborates a transit loss or damage for this line. Cite BOL vs Received.
internal the BOL states the full PO quantity was shipped and receiving recorded less than that, but nothing corroborates a transit loss for THIS line -- either there is no carrier exception on file at all, or the exception on file names a DIFFERENT SKU than this one. It looks like a transit-loss pattern, but nothing on file backs it for this line; the likely explanation is a dock miscount, not a proven transit loss. Cite BOL vs Received.
insufficient_evidence the three quantities disagree in a shape none of the other three calls resolve cleanly -- most often the BOL and receiving AGREE with each other but both disagree with the PO (an apparent over-shipment, not a shortfall), or more than one of the three quantities is off at once. There is no clean documentary basis for a liability call. Cite PO vs Received.
⚠︎ A CARRIER EXCEPTION EXISTING ON THE CASE IS NOT, BY ITSELF, EVIDENCE THAT THE CARRIER IS LIABLE FOR THIS LINE. Check whether the exception names the SAME SKU as the discrepant line. An exception about a different SKU is a real note about a different problem -- it does not excuse this one, and calling the carrier liable on its strength alone is exactly the mistake this task exists to catch.
The notice you draft must name the discrepant SKU and state which two of the three quantities (po, bol, received) disagree, their values, and the delta between them -- e.g. "BOL states 240 units shipped; receiving recorded 210 units, a shortfall of 30." Use whichever pair the liable_party value above says to cite. delta is the absolute difference between qty_a and qty_b, never signed.
You never issue a chargeback or a credit yourself, and you never make the final liability determination -- you propose a probable liable party and its documentary basis for receiving/AP to confirm and act on.
Build the case file for the case below and propose a probable liable party.
Return a JSON object with exactly these keys: "liable_party" (one of: vendor, carrier, internal, insufficient_evidence), "discrepant_sku" (the SKU of the one discrepant line), "doc_a" and "doc_b" (the two of "po", "bol", "received" that disagree, per the liable_party value's citation rule above), "qty_a" and "qty_b" (their quantities), "delta" (the absolute difference between qty_a and qty_b), "notice_text" (the drafted notice, one or two sentences).
CASE CASE-00001 (PO PO-356479, carrier Union Point Freight)
Lines:
SKU-4001 PO 334 BOL 334 Received 334
SKU-4002 PO 105 BOL 124 Received 124
Carrier exception: none on file
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"liable_party": "insufficient_evidence",
"discrepant_sku": "SKU-4002",
"doc_a": "po",
"doc_b": "received",
"qty_a": 105,
"qty_b": 124,
"delta": 19,
"notice_text": "PO ordered 105 units of SKU-4002; receiving recorded 124 units, an overage of 19. BOL also states 124 units, so this appears to be an over-shipment rather than a transit loss."
}
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check who's liable when a retailer's delivery doesn't add up — 36 cases. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: exact match on liable_party (the core 4-way classification); false_confident_call (a confident, non-insufficient_evidence call on a gold insufficient_evidence case) counted over exactly that denominator; wrong_sku_exception_misread (a trap case credited to the carrier anyway) counted over exactly the corpus's planted trap denominator; and discrepant_sku plus quantity-citation accuracy checked separately, requiring the same document pair and delta, not just two correct numbers. No judge model. The same function scores both evals/baseline.py's free floor and evals/run.py's real run.
36cases
36source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED36 / 36liable accuracy pct — cases with a parsed liable-party answerDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED36 / 36answered pct — cases askedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 6false confident call rate pct — gold insufficient_evidence casesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 6wrong sku exception misread rate pct — planted trap casesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED36 / 36discrepant sku accuracy pct — cases scoredDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED36 / 36quantity citation accuracy pct — cases scoredDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/scoring.py) is pure code, exact match against a mechanically-derived gold determination -- there is no judgement to validate, only string/number comparison. What WAS validated: tools/build_corpus.py's gold liable_party, discrepant_sku, doc_a/doc_b/qty_a/qty_b/delta and trap flag are all RECOMPUTED by classify_case() from the same generated case record the model reads, never carried over from the target category the corpus builder aimed for -- and --verify independently re-derives every row from data/cases.jsonl on disk and asserts zero drift, confirmed this session (36 cases checked, 0 drift; 6 trap cases of 14 internal, all resolving as documented).
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One case
1,000 cases
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.002141
$2.14
23%
Same work, 1× the bill
The same cases, the same tokens — only the rate card changed. And on that card about 23% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether reasoning ('thinking') is left on or explicitly disabled for this model -- the app's own /api/check always disables it; the registered run left it at provider default. The two configurations have not been priced against each other on this corpus.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page. Reasoning was also left at provider default (on) for this run -- see Business.not_good_enough -- so even the projected figure prices tokens the shipped app's own reasoning-off configuration would not spend.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing -- exact match on liable_party/SKU/doc-pair/delta is pure code. A measured $0.00, not an unpriced one.
The gradersOne way to grade, and why it is the only one
The floor blames the carrier on ANY exception note present on the case, without checking which SKU it names -- so it misclassifies all 6 planted trap cases as carrier when gold says internal (a 100% wrong_sku_exception_misread rate, 6 of 6), while getting every other case right: the corpus never plants an exception note on a vendor or insufficient_evidence case, and the 8 internal_clean cases carry no exception at all, so presence-of-exception and correct-SKU-match agree everywhere except the 6 planted traps. Its overall liable_party accuracy still reaches 83.3% (30 of 36) for exactly that reason. The fast tier's 0% wrong_sku_exception_misread and 100% liable_party accuracy on the identical 36 cases is the real, measured gap a model closes that this floor cannot.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Liable-party exact match, plus false-confident-call, wrong-SKU-trap and quantity-citation checks Does the model's liable_party match gold's four-way classification? Does it avoid a confident call on a case gold says has no clean documentary basis (false_confident_call)? Does it avoid crediting a carrier exception note that names a different SKU than the discrepant line (wrong_sku_exception_misread)? Do the cited discrepant SKU, document pair, quantities and delta match gold?
$0.00
no
yes
the fast tier 100.0% liable accuracy · 1 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, and cleanly on the one axis this corpus was built to test: the fast tier and the free exception-first floor tie on the 30 non-trap cases (both agree when exception-presence and correct-SKU-match point the same way), then diverge sharply on the 6 planted trap cases -- 0% wrong_sku_exception_misread (fast tier) vs 100% (floor, i.e. the floor gets every one wrong) -- because the floor's presence-only check answers exactly what the trap is built to make it answer, and the fast tier does not. A grader that could not tell the two apart would not produce a gap that lines up this precisely with the corpus's own documented trap count.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Resolving a receiving-discrepancy case's liable party from its PO/BOL/receiving quantities and any carrier exception note
the fast tier, over the free exception-first floor
100% liable_party accuracy and 0% wrong_sku_exception_misread this run, against the floor's 83.3% and 100% wrong_sku_exception_misread rate on the exact cases built to be misread on exception presence alone.
the exception-first floor for anything beyond the non-trap cases, where blaming the carrier on any exception happens to agree with the correct call -- it is not a competitor, it is the honest floor a model has to clear.
Deciding whether a case has a clean documentary basis for a liability call at all
the fast tier's liable_party verdict, read alongside whether it answered insufficient_evidence rather than forcing a guess
0 of 6 gold insufficient_evidence cases this run were given a confident vendor/carrier/internal call (false_confident_call=0) -- the expensive-direction error this kit's guardrail exists to catch.
trusting a confident verdict without checking whether reasoning was left on for the call that produced it -- see Business.not_good_enough; this run's own configuration has not been re-measured with reasoning explicitly disabled the way the live app runs it.
What this kit is not
a proposed liable party and its documentary basis for receiving/AP to review -- never a chargeback, never a credit, never a final determination
src/resolve.py and src/app.py have no function anywhere that issues a chargeback, posts a credit, or records a final liability decision.
issuing a chargeback or credit straight off this kit's drafted notice with no receiving/AP confirmation step.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
vendor
BOL states less than the PO ordered; receiving matches the BOL
8
CASE-00006: PO ordered 86 units of SKU-4016, BOL states only 78 shipped, receiving recorded 78 (matching the BOL) -- a shortfall of 8. Gold and model both: vendor.
carrier
Full PO qty shipped, less arrived, AND the carrier's own exception note names THIS SKU
8
CASE-00002: BOL states 172 units of SKU-4004 shipped, receiving recorded 132 -- a shortfall of 40. The carrier exception on file names SKU-4004 directly, corroborating a transit loss for this exact line. Gold and model both: carrier.
internal_clean
Same shortfall shape as carrier, but no exception note is on file at all
8
CASE-00011: BOL states 340 units of SKU-4028 shipped, receiving recorded 243 -- a shortfall of 97. No carrier exception is on file for this case. Gold and model both: internal.
internal_trap_resolved
Same shortfall shape as carrier, but the exception note on file names a DIFFERENT SKU -- resolved correctly
6
CASE-00010: BOL states 169 units of SKU-4026 shipped, receiving recorded 129 -- a shortfall of 40. A carrier exception IS on file, but it names SKU-4027, a different line. Gold internal (a trap case). The exception-first floor blames the carrier on any…
insufficient_evidence
BOL and receiving agree with each other but not the PO (an apparent over-shipment), or more than one quantity disagrees at once
6
CASE-00001: PO ordered 105 units of SKU-4002; BOL and receiving both state 124 -- an apparent over-shipment, not a shortfall, with no clean documentary basis for a vendor/carrier/internal call. Gold and model both: insufficient_evidence.
What we could NOT verify
Whether this run's 100% liable_party accuracy and 0% false_confident_call/wrong_sku_exception_misread are stable properties of this model at these settings, or a one-off -- one recorded run is not a history.
Whether the same figures hold with reasoning explicitly disabled. The registered run (r001-rcv-disc) left provider-side reasoning at its documented default (on); the shipped app's own /api/check hardcodes it off. No run exists at the app's actual setting -- see Business.not_good_enough.
Whether resistance holds against an adversarial or fabricated carrier exception note -- no red-team run exists for this kit (see redteam_page).
Whether a real receiving operation's own genuinely missing document (rather than a document that resolves unclearly) would be detected as incomplete rather than misread as a clean or exception case -- see Data.breaks_on.
Whether a real carrier's own exception-note phrasing (rather than this corpus's six templates) would reproduce these figures -- see Data.bring_your_own_boundary.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
997.42
547.39
4,090 ms
$0.002141
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The exception-first baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 35907 input / 19706 output tokens priced at Google Gemini 3 Flash's published rate, the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
Every SKU line on the case (PO, BOL and receiving quantities) plus the carrier exception note or its absence -- fixed size per call by construction (input tokens ranged 971-1024 across all 36 calls, essentially flat), the floor every call pays regardless of which liable party governs.
Reasoning ('thinking') was left at the provider's default (on) rather than explicitly disabled for this run -- it consumed 76.9% of the average output-token budget (421.1 of 547.4 tokens/call) and is already priced into cost_per_query_usd, since a provider bills reasoning tokens as completion tokens. src/app.py's live UI hardcodes thinking off, a different, unmeasured configuration -- see Business.not_good_enough.
The fixed system prompt (3,117 characters -- the four-value taxonomy, citation vocabulary and the wrong-SKU-exception warning) is sent in full on every call, the largest single fixed cost per case.
Your volumeWhat it costs at your volume
Linear in cases: each call is independent and self-contained, with no shared context or retrieval step to amortise. This run's 36 cases cost about $0.0771 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $0.7710 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volumereceiving-discrepancy cases resolved per day/week -- this run resolved 36 in one pass
What it replacesreceiving/AP staff reading a case's PO/BOL/receiving quantities and any carrier exception note by hand and proposing a liable party grounded in which SKU the exception actually names
Time saved per itemnot measured here -- depends on how long a manual case-file review takes at the reader's own company
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run -- unlike data-reconcile's fast/reasoning comparison -- so this page prices one model, not a trade-off; see Eval.could_not_verify for what a second run would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
35,907input tokens · this run
19,706output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 36 receiving-discrepancy cases resolved, 36 cases scored by pure code. This run (r001-rcv-disc) answered 36 of 36 liable-party calls -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.031
$0.031
$0.86
2026-09-12
gemini-3-flash
Google
$0.077
$0.077
$2.14
2026-09-18
gemini-3-8-flash
Google
$0.101
$0.101
$2.80
2026-09-18
llama-5
Meta
$0.129
$0.129
$3.57
2026-09-18
claude-haiku-4-5
Anthropic
$0.134
$0.134
$3.73
2026-09-12
grok-4-5
xAI
$0.190
$0.190
$5.28
2026-09-18
grok-4-6
xAI
$0.190
$0.190
$5.28
2026-09-18
claude-sonnet-5
Anthropic
$0.269
$0.269
$7.47
2026-09-12
gemini-3-1-pro
Google
$0.308
$0.308
$8.56
2026-09-18
gpt-5-6-terra
OpenAI
$0.308
$0.308
$8.56
2026-09-12
gpt-5-6-sol
OpenAI
$0.538
$0.538
$14.94
2026-09-12
claude-opus-4-8
Anthropic
$0.672
$0.672
$18.67
2026-09-12
claude-opus-5
Anthropic
$0.672
$0.672
$18.67
2026-09-12
claude-fable-5
Anthropic
$1.344
$1.344
$37.34
2026-09-18
claude-fable-5-1
Anthropic
$1.344
$1.344
$37.34
2026-09-18
gpt-6-astra
OpenAI
$1.344
$1.344
$37.34
2026-09-17
Read this against the numbers above
REASONING WAS NOT DISABLED FOR THIS WORKLOAD. Every row below prices THIS run's own token counts, which include reasoning tokens the live app's own /api/check never generates (it hardcodes thinking off). A forker's real bill on other models depends on whether that model has an equivalent reasoning toggle and whether it is left on.
NO QUALITY IS IMPLIED. Only the model that produced Eval.scores has been scored against this corpus -- every row here is a price, not a recommendation.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 36 receiving-discrepancy cases from a fixed seed (SEED=20260819) -- 2-4 SKU lines per case, one genuinely discrepant, plus an optional carrier exception note naming exactly one SKU. Gold (liable_party, discrepant_sku, doc_a/qty_a, doc_b/qty_b, delta, trap) is RECOMPUTED from the actual generated case record by classify_case(), never from the target category the generator aimed for -- --verify independently re-checks every row against the record actually written to disk.
You change it to: Point it at your own SKUs, carriers and case shapes. src/resolve.py and src/prompt.py read a case by its own fields (lines[].sku/po_qty/bol_qty/received_qty, carrier_exception) and do not care where the numbers came from.
tools/build_corpus.py
# Generate the receiving-discrepancy cases this kit checks against, from a fixed seed.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
DOCS = ("po", "bol", "received")
CASE_PLAN = (
CARRIERS = [
EXCEPTION_TEMPLATES = [
def classify_case(lines, carrier_exception):
def _clean_line(sku, rng):
src/prompt.pythe prompt
The four-value liability taxonomy and the citation vocabulary, declared once and read from here by build() and parse(). States the wrong-SKU-exception failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover.
src/prompt.py
# Assemble the one prompt this kit sends, and parse the one reply it gets back.
LIABLE_PARTIES = ("vendor", "carrier", "internal", "insufficient_evidence")
DOCS = ("po", "bol", "received")
LIABLE_MEANINGS = {
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _case_block(case):
def build(case, prompt=DEFAULT_PROMPT):
def parse(raw):
src/resolve.pythe AI layer
Loads one case by id, calls the model once with every SKU line and the carrier exception note (or its absence), parses the reply into a liable party, discrepant SKU, cited quantity pair, delta and drafted notice. The whole AI layer, deliberately short -- the same split data-reconcile's reconcile.py and fin-payrun's payrun.py both make. Never issues a chargeback or credit and never makes the final liability call -- there is no function here or in src/app.py that does either.
src/resolve.py
# Build one receiving-discrepancy case file and propose a probable liable party: load the case,
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CASES = os.path.join(HERE, "data", "cases.jsonl")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
MAX_TOKENS = 3000
def cases():
def load_gold():
def check(cfg, case, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pyspend control
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
src/app.pythe app
Static files and three JSON endpoints (/api/state, /api/case, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-rcv-disc) actually used; see Business.not_good_enough.
src/app.py
# The minimal local UI. Standard library only -- python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8790"))
class H(BaseHTTPRequestHandler):
def main():
evals/scoring.pythe scorer
Four-way liable_party exact match, pure code: false_confident_call (a confident call on a gold insufficient_evidence case) and wrong_sku_exception_misread (a trap case credited to the carrier anyway) reported as their own named figures rather than folded into a generic 'wrong' bucket, plus discrepant_sku and quantity-citation accuracy checked separately.
evals/scoring.py
# Score a set of predicted liable-party determinations against gold. Pure code, shared by
LIABLE_PARTIES = ("vendor", "carrier", "internal", "insufficient_evidence")
def _pct(n, d):
def score(records, gold):
evals/baseline.pyfree baseline
A fixed rule, written the way a person free-texting a quick classifier would write one, that blames the carrier on ANY exception note without checking which SKU it names -- so it misclassifies every planted trap case by construction. The honest, narrow floor a model has to clear.
evals/baseline.py
# What a simple, dumb rule catches, over the same corpus. Free. No key, no dependency, no model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def classify(case):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 36 receiving-discrepancy cases from a fixed seed (SEED=20260819) -- 2-4 SKU lines per case, one genuinely discrepant, plus an optional carrier exception note naming exactly one SKU. Gold (liable_party, discrepant_sku, doc_a/qty_a, doc_b/qty_b, delta, trap) is RECOMPUTED from the actual generated case record by classify_case(), never from the target category the generator aimed for -- --verify independently re-checks every row against the record actually written to disk. A swap seam.
src/prompt.pyThe four-value liability taxonomy and the citation vocabulary, declared once and read from here by build() and parse(). States the wrong-SKU-exception failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover.
src/resolve.pyLoads one case by id, calls the model once with every SKU line and the carrier exception note (or its absence), parses the reply into a liable party, discrepant SKU, cited quantity pair, delta and drafted notice. The whole AI layer, deliberately short -- the same split data-reconcile's reconcile.py and fin-payrun's payrun.py both make. Never issues a chargeback or credit and never makes the final liability call -- there is no function here or in src/app.py that does either.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from. A swap seam.
src/budget.pyAn append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
evals/scoring.pyFour-way liable_party exact match, pure code: false_confident_call (a confident call on a gold insufficient_evidence case) and wrong_sku_exception_misread (a trap case credited to the carrier anyway) reported as their own named figures rather than folded into a generic 'wrong' bucket, plus discrepant_sku and quantity-citation accuracy checked separately.
evals/baseline.pyA fixed rule, written the way a person free-texting a quick classifier would write one, that blames the carrier on ANY exception note without checking which SKU it names -- so it misclassifies every planted trap case by construction. The honest, narrow floor a model has to clear.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 997 input and 547 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's cases are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any text the model actually read this run. In a real deployment the carrier exception note would arrive from an external carrier's own EDI feed or claims portal message -- exactly the kind of externally-authored input data-reconcile's agreement-poisoning attack measures for a different artifact, and this kit's architecture treats it as trusted input with no verification step. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and one boundary that should exist does not
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every carrier exception note is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's exception note would carry externally-authored text (see posture), and that surface has not been attacked. The gates above are boundaries confirmed by reading the code, not payloads run through it; one is a negative result, not a guarantee. Confirmed by reading the code, not by a run, on 2026-08-20 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a determination ever issue a chargeback or post a credit?
A drafted liable-party call or notice could plausibly trigger a chargeback or a credit posting.
No code path does. src/resolve.py::check() and src/app.py's /api/check both return an answer only; neither writes to data/cases.jsonl or any other store -- confirmed by reading every call site.
Can the model's own liable_party and its cited doc_a/doc_b/delta disagree with each other, or with the record, and still reach the reader unflagged?
One would expect a liable-party call that contradicts its own cited quantities to be caught before it renders.
NO -- this is a real gap, not a guarantee. Nothing in src/resolve.py or src/app.py cross-checks the model's own doc_a/qty_a/doc_b/qty_b/delta against the case record, or re-derives liable_party from them, before the UI shows them; the rule lives only in the system prompt. See Guardrails.fails.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/check handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Can a carrier exception note naming the wrong SKU still be credited to the carrier?
A model reading 'an exception exists' without checking which SKU it names could call the carrier liable on a case where the exception is unrelated.
Measured, not just guaranteed: 0 of 6 planted trap cases this run were misread this way (wrong_sku_exception_misread=0) -- but this is a property of THIS run's replies, not a code-level check; nothing recomputes the SKU match before rendering. See Guardrails.fails.
Each boundary above was checked by reading the call sites, not by an attack trial. Gate 2 and gate 4's underlying code path are the ones that do NOT hold structurally -- see Guardrails.fails for the same findings from the enforcement side.
The result0 attack trials, and one boundary that does NOT hold in code (only in the prompt): the model's own cited doc_a/doc_b/qty_a/qty_b/delta and its liable_party call are never cross-checked against the case record or against each other before the UI renders them. Two boundaries (no write path, key redaction) hold, confirmed by reading the code.
1externally-authored field a live deployment would carry (the carrier exception note) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no case's carrier exception note was authored by an outside party. A real deployment's exception note would be exactly the kind of externally-supplied text data-reconcile's agreement-poisoning red-team run measured being poisoned for a different artifact; that question is open for this kit.
Read this twice
This kit has no code-level check on its own model's liable_party/citation consistency. Everything from the wrong-SKU-exception rule to the citation-quantity grounding lives in the system prompt alone -- src/resolve.py and src/app.py pass the model's parsed reply straight through. On this run every field happened to be internally consistent and every trap case was correctly resolved (0 wrong_sku_exception_misread), but that is a property of THIS run's replies, not a guarantee the code provides. A future version should recompute whether the cited doc_a/doc_b/delta actually matches the case record, and whether an exception's named SKU actually equals the discrepant SKU, before trusting a live determination.
HonestyWhat this does not prove
Whether an injected instruction inside a carrier exception note (e.g. 'this exception covers the whole shipment') could move a liable-party call -- no red-team run exists for this kit.
Whether the live app's hardcoded thinking=off setting changes resistance to a hostile exception note relative to this run's provider-default-on configuration -- untested either way.
Whether a code-level consistency check (Guardrails.add_first) would catch a real liable_party/citation disagreement in practice -- none has been built or exercised.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
A carrier exception note is only evidence for the SKU it actually names, never for the case in general -- a case is never resolved to carrier liability on an exception's presence alone.
src/prompt.py -- SYSTEM, stated as an explicit boundary the model must respect. Prompt-level ONLY: unlike data-reconcile's live citation-fidelity check, no code in src/resolve.py or src/app.py cross-checks the model's own cited doc_a/qty_a/doc_b/qty_b/delta against the case record, or re-derives liable_party from the exception's named SKU, before the UI renders them.
EvidenceDoes it hold?
What
Measured
No code path in this kit issues a chargeback or posts a credit
0 of 36 calls in r001-rcv-disc resulted in any write to data/cases.jsonl or any other store -- src/resolve.py::check() and src/app.py's /api/check both only ever return an answer.
The prompt-level wrong-SKU rule itself held on this run
0 of 6 gold trap cases in r001-rcv-disc were credited to the carrier on the strength of an exception naming a different SKU (wrong_sku_exception_misread=0) -- but see fails below for what does, and does not, enforce this.
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK. The rule that an exception is only evidence for the SKU it names lives only in src/prompt.py's SYSTEM text -- nothing in src/resolve.py or src/app.py recomputes it. See fails above.
It does not verify the cited doc_a/doc_b/qty_a/qty_b/delta against the case record mathematically -- see fails.
It does not make the liable-party classification itself correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against a hostile or fabricated carrier exception note -- no red-team run exists for this kit (see the security page).
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 17 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
16 measured by the latest run1 need the model half
Metric
Owner
Role
Why this one
liable-party-exact-match-plus-trap-and-citation
Liable-party exact match, plus false-confident-call, wrong-SKU-trap and quantity-citation checks
alarm
wrong_sku_exception_misread as a raw count, never folded into liable_accuracy -- it is the expensive-direction error this kit's guardrail exists to catch; the per-liable-party breakdown specifically, since a model that credits any exception present fails exactly on the trap subset of internal cases (see baseline_note); answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed — alarm on Any nonzero wrong_sku_exception_misread on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
36
different corpus — nothing is comparable
corpus.bytes
13,434
cases edited — the count held, the bytes did not
split.count
36
the characters per case's assembled block count moved — a different set was scored
split.size_p50
283
the median size of one characters per case's assembled block moved
split.size_p95
301
the 95th-percentile size of one characters per case's assembled block moved
dataset.rows
36
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
liable-party accuracy
not yet known
36 answered cases
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat -- r001-rcv-disc ran once.
false confident call
0 -- zero this run, this kit's headline guardrail metric
6 gold insufficient_evidence cases
measured, one run
wrong-SKU exception misread
0 -- zero this run, this kit's second guardrail metric
measured, one run -- the per-party breakdown of the overall liable-party accuracy figure above; not yet known to be stable across a repeat.
cases answered
not yet known
36 cases
100% this run -- every case drew a liable-party classification, none silently dropped. No repeat -- r001-rcv-disc ran once.
reasoning-token share of output
not yet known
36 calls
76.9% this run (15161 of 19706 output tokens), one run, one setting -- reasoning left at provider default, never compared against the live app's thinking-off setting. Narrative only, not a registered metric key -- the run record does not carry a reasoning-token-share field, only raw input/output totals.
latency
not yet known
36 calls
p50 4090 ms, p95 9869 ms on r001-rcv-disc -- reasoning left on. One recorded run, not a distribution.
input volume
0 -- fixed by the corpus and the prompt, not the model
36 calls
35907 input tokens on r001-rcv-disc. Any movement means the prompt or the corpus changed.
output volume
not yet known
36 calls
19706 output tokens on r001-rcv-disc -- model-specific, and includes whatever reasoning the provider chose to spend.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-rcv-disc 2026-08-20
answered, %
100.0
carrier accuracy, %
100.0
discrepant sku accuracy, %
100.0
false confident call
0
false confident call rate, %
0.0
input tokens, whole run
35907
insufficient evidence accuracy, %
100.0
internal accuracy, %
100.0
model latency p50 ms
4090.00
model latency p95 ms
9869.00
liable accuracy, %
100.0
output tokens, whole run
19706
quantity citation accuracy, %
100.0
vendor accuracy, %
100.0
wrong sku exception misread
0
wrong sku exception misread rate, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 16 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether reasoning is left on or explicitly disabled
not measured -- this run left it on (provider default); the live app always disables it; the two have never been compared on this corpus
reasoning
src/app.py hardcodes thinking=THINKING_OFF; r001-rcv-disc's own top-level thinking field is null (provider default), while every one of its 36 per-case records carries a nonzero reasoning_tokens (154-1,206). See Business.not_good_enough.
the planted trap (a carrier exception naming a different SKU, 6 of 14 internal cases in tools/build_corpus.py)
wrong_sku_exception_misread on the 6 trap cases: 100% (exception-first floor, all 6 misread) -> 0% (the fast tier) on the identical cases
measured
results/eval-b000-exception-first.json vs results/eval-r001-rcv-disc.json, same 36 cases, one variable (which classifier reads the record).
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
liable-party accuracy
nothing yet
false confident call
any nonzero value, on any run
wrong-SKU exception misread
any nonzero value, on any run
discrepant SKU accuracy
any drop below 100% on a re-run
quantity citation accuracy
any drop below 100% on a re-run
per-liable-party accuracy
any liable party falling below 100% on a re-run
cases answered
nothing yet
reasoning-token share of output
nothing yet
latency
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
nothing yet
NextThe three you would add first
A live, code-level consistency check: confirm the model's cited doc_a/qty_a/doc_b/qty_b actually match the case record, and delta equals their absolute difference, before the UI renders them.Currently nothing catches a model that reports quantities disagreeing with its own case record -- the same 'verdict disagrees with its own extracted values' shape as data-reconcile's false-clean cases, and this kit has no check for it at all, live or offline, beyond the eval's own comparison to gold.
A live check that a liable_party='carrier' call's exception note actually contains the discrepant SKU string.The system prompt states the rule; nothing in code enforces it before a reader sees it. See fails.
A red-team run against the carrier exception note text, the way data-reconcile attacked its supplier agreement.In a real deployment the exception note arrives from an external carrier this kit's architecture treats as trusted with no verification step -- unmeasured for this kit (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting in src/resolve.py.
What this cannot tell you
Whether the same 100%/0% figures hold with reasoning explicitly disabled, matching the live app's own setting -- see Business.not_good_enough.
Whether a live, code-level consistency check (see add_first) would ever have caught a real disagreement -- this run's own model never produced one to test against.
Whether the prompt-only wrong-SKU and citation-grounding rules hold against a hostile or fabricated carrier exception note -- no red-team run exists for this kit.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is a comment-only file explaining that the emptiness is load-bearing. The whole resolving decision is three files: src/prompt.py, src/resolve.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
36 receiving-discrepancy cases, generated from a fixed seed, never fetched. Recomputing gold from the same generated case record (never the target category that seeded it) via classify_case() is what keeps the internal-consistency check honest -- see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the four-value liability taxonomy and the citation vocabulary are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. Passing thinking through is a kwarg on one call site, not a client-library upgrade.
evaluation
evals/scoring.py
eval harnesses
exact match over a four-value liable-party vocabulary plus the two named guardrail metrics is a dict comprehension and a couple of counters, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per case -- load, assemble, prompt, call, parse -- with no branching and no state carried between cases. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A fifth liable-party value this kit's fixed four-value vocabulary doesn't name needs a hand-written LIABLE_PARTIES/LIABLE_MEANINGS entry plus new corpus category logic in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here.
No built-in output validation beyond src/prompt.py::parse()'s own tolerant-but-not-creative JSON parse -- a framework with a schema-and-consistency-check layer baked in might catch a liable_party/citation disagreement before it renders; this kit's stdlib-only design did not build one (see Guardrails.fails).
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-rcv-disc on the fast tier, 2026-08-20. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
4,090 ms
not yet known
nothing yet
Model, p95
9,869 ms
not yet known
nothing yet
Input tokens
35,907
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
19,706
not yet known
nothing yet
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the rulebook is data read whole into the prompt.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-20, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
receiving-discrepancy cases
data/cases.jsonl -- 36 cases, generated once from a fixed seed by tools/build_corpus.py; nothing else knows what a case's lines and exception note are
one whole case -- every SKU line and the carrier exception note (or its absence) -- sent in the one call; src/app.py and src/resolve.py both read it by case_id, never modified after generation
gold determinations
data/gold.jsonl -- 36 rows, computed by tools/build_corpus.py's classify_case() by re-reading the same generated case record, never carried over from the target category that seeded it
never -- evals/scoring.py is pure code, no model, no key. src/resolve.py::load_gold()'s own docstring: NEVER read by check().
the key
.env -- never committed
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
rulebook
tools/build_corpus.py regenerates the whole corpus -- 36 cases -- byte-identically from a fixed seed (SEED=20260819) every time it is run, with an independent --verify pass re-deriving every gold row from the case records actually written to disk.
13434 bytes, one file (data/cases.jsonl), 36 cases -- see Data.corpus. --verify confirmed 0 drift across all 36 rows this session. (lenses.Data.corpus, tools/build_corpus.py)
a real receiving operation's case backlog is continuous, not a fixed seed, and a case with a genuinely missing document (rather than an unclear one) cannot be recognised as such by this kit at all (see Data.breaks_on) -- neither was measured
point tools/build_corpus.py at your own SKUs and exception phrasing, and every published accuracy/wrong-SKU figure is void -- they are this corpus's own phrasing templates and trap design (see Data.bring_your_own_boundary), not a property of the model
model
The four-value liability taxonomy (vendor, carrier, internal, insufficient_evidence) and the wrong-SKU-exception warning are declared once in src/prompt.py and sent in full inside the system prompt on every call -- a FIXED rulebook that does not vary by case. One call per CASE carries every SKU line and the carrier exception note (or its absence) alongside it, behind src/adapters/__init__.py, reasoning left at the provider's default (on) -- the configuration r001-rcv-disc ships, and a DIFFERENT configuration from the one src/app.py's live UI actually runs (thinking explicitly off).
The rulebook portion of the system prompt is 3,117 characters, sent identically on all 36 calls in r001-rcv-disc -- see LLM.prompt_verbatim for the exact text. 36 of 36 cases in r001-rcv-disc returned a reply that parsed cleanly (0 failures, finish_reason 'stop' on all 36), with the largest single call using 1,322 of the 3,000-token MAX_TOKENS budget (1,206 of it reasoning) -- real headroom. (lenses.Business.not_good_enough, results/eval-r001-rcv-disc.json, src/resolve.py)
a fifth liable-party value this kit's fixed four-value vocabulary doesn't name needs a hand-written LIABLE_PARTIES/LIABLE_MEANINGS entry plus new corpus category logic in tools/build_corpus.py -- never learned from data alone, and never measured here. Separately, a call needing meaningfully more reasoning than this run's worst case (1,206 tokens) was never measured against MAX_TOKENS=3000.
a different liable-party set (added, removed or redefined values) invalidates the whole eval at once -- gold and grading in evals/scoring.py are both keyed to this exact four-value vocabulary. Verdicts are also per-model and this configuration ran once -- the live app's own thinking-off setting has never been scored against this corpus; see Business.not_good_enough.
labels
data/gold.jsonl, 36 rows -- liable_party, discrepant_sku, doc_a/qty_a, doc_b/qty_b, delta and the trap flag are all RECOMPUTED from the same generated case record the model reads by classify_case(), never from the target category that seeded the generator, and independently re-derived by --verify against the case records actually written to disk.
your own receiving cases: hand-label the gold, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-labelled gold has an error rate this kit has never measured
liable-party and wrong-SKU-exception figures over this set reflect ONE planted trap family (a carrier exception naming a different SKU) and ONE assumption (every case's documentary evidence is complete) -- a real receiving operation's other failure shapes (see Data.breaks_on) are untested
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
liable_party is carrier, but the cited carrier exception note does not actually name the discrepant SKU
nothing in code would catch this -- the rule is prompt-only (see Guardrails.fails); this run never produced one, but the code does not prevent it
re-read the carrier_exception text yourself and confirm it names the discrepant_sku before trusting a carrier verdict, until Guardrails.add_first's live check exists (lenses Guardrails.fails, src/resolve.py, src/app.py)
the cited doc_a/qty_a/doc_b/qty_b do not match the case's own lines, or delta does not equal their absolute difference
the model stated citation fields inconsistent with its own case record; nothing in code re-derives them before rendering -- this run never produced this disagreement (see Eval.scores) but the code does not guard against it
recompute the cited quantities against the case's own lines yourself before trusting the fields as printed (lenses.Guardrails.fails, src/prompt.py)
No machine symptom — this failure leaves no trace in any output.
reasoning ('thinking') left at provider default for the registered run, while the live app hardcodes it off -- a reader who re-runs this kit's own app will see different latency and cost than this report publishes, and nobody has measured whether accuracy differs too. See Business.not_good_enough.
Concurrency and GPU sizing -- one serial call per case, nothing measured past 36. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- Cost prices every call at the peak, cache-miss rate for exactly this reason. Whether a real carrier's own exception-note phrasing or a real receiving operation's own missing-document shape would reproduce these figures -- see Data.bring_your_own_boundary. Whether reasoning explicitly disabled (the live app's own setting) changes any figure on this page -- every scored run here left it at provider default. Whether a hostile or fabricated carrier exception note could move a liable-party call -- no red-team run exists for this kit.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Liable-party exact match, plus false-confident-call, wrong-SKU-trap and quantity-citation checks
Check who's liable when a retailer's delivery doesn't add up
PresenterOpens the private repo. Visible to admins only.
In one lineLiable-party exact match, plus false-confident-call, wrong-SKU-trap and quantity-citation checks
Does the model's liable_party match gold's four-way classification? Does it avoid a confident call on a case gold says has no clean documentary basis (false_confident_call)? Does it avoid crediting a carrier exception note that names a different SKU than the discrepant line (wrong_sku_exception_misread)? Do the cited discrepant SKU, document pair, quantities and delta match gold?
$0.00per 1,000 cases
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Case
CASE-00010
Carrier
Union Point Freight
Carrier exception on file
pallet found wet on arrival, partial refusal noted for SKU-4027
Gold liable party
internal
Exception-first floor's call
carrier
The model's call
internal
The model's drafted notice
Discrepant SKU SKU-4026: BOL states 169 units shipped; receiving recorded 129 units, a shortfall of 40. The carrier exception names SKU-4027, not SKU-4026, so there is no documented transit-loss basis for this line; it is likely a dock miscount.
Scored as
correct -- the trap case the free exception-first floor fails (a wrong_sku_exception_misread) and the fast tier does not
Grader
Verdict
Why
Liable-party exact match, plus false-confident-call, wrong-SKU-trap and quantity-citation checks
correct
CASE-00010: BOL states 169 units of SKU-4026 shipped; receiving recorded 129, a shortfall of 40. A carrier exception IS on file for this case ('pallet found wet on arrival, partial refusal noted for SKU-4027'), but it names SKU-4027 -- a different line -- not the discrepant SKU-4026. Gold liable_party internal (a trap case: exception present, wrong SKU). The exception-first floor blames the carrier on any exception present and answers carrier -- a wrong_sku_exception_misread. The fast tier answered internal, correctly declining to credit an exception naming a different SKU.
The formulaWhat it computes
liable_accuracy = correct liable_party / answered cases, over answered=36. false_confident_call = a non-null, non-insufficient_evidence call on a gold insufficient_evidence case, counted over exactly that denominator (6). wrong_sku_exception_misread = a call of 'carrier' on a gold trap case (exception on file naming a different SKU), counted over exactly the planted-trap denominator (6). discrepant_sku_accuracy and quantity_citation_accuracy are their own exact-match rates, the latter requiring the same (doc, qty) pair as an unordered set plus a matching delta.
The analysisWhat it actually did
Model
Result
the fast tier
100.0% liable accuracy · 1 more measured on this row
In operationWhat to monitor
Reference standard: tools/build_corpus.py's classify_case(), which reads the same case record IN ORDER and computes the gold liable_party, discrepant_sku, doc_a/qty_a/doc_b/qty_b, delta and trap flag -- re-derived independently via --verify against the case records actually written to disk, asserted zero drift.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own planted trap design, which comes from tools/build_corpus.py's fixed templates.
Watch these
wrong_sku_exception_misread as a raw count, never folded into liable_accuracy -- it is the expensive-direction error this kit's guardrail exists to catch
the per-liable-party breakdown specifically, since a model that credits any exception present fails exactly on the trap subset of internal cases (see baseline_note)
answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed
Alarm on
Any nonzero wrong_sku_exception_misread on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
How tight can the band be? There is no tolerance band here -- every field is an exact match (liable_party, SKU string, doc labels, numeric delta); the liability-resolution task has no continuous quantity to round.
Cadence: Re-run on any change to src/prompt.py, tools/build_corpus.py, or MAX_TOKENS in src/resolve.py -- the first changes what is asked, the second changes what is asked ABOUT, the third bounds how many verdicts can come back at all.
The decisionWhen to reach for it
Use it
The gold determination is DERIVED from the same generated case record the model reads -- true of every kit corpus, never true of a real deployment's own receiving-discrepancy trail.
Do not use it
The truth is not known in advance -- the normal state of a real receiving dock, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning