Find why a store's stock count doesn't match the system
Every cycle count throws off gaps, and each one means digging through the item's receiving, transfer and scan history. This app reads that history, drafts the likely cause with the exact line behind it, and leaves the decision to inventory control.
PresenterOpens the private repo. Visible to admins only.
For inventory controlCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
An inventory-control analyst at a retail chain or warehouse, checking the gaps each cycle count throws off.
✕Today's manual process
1Open each gap from the count report, one item and location at a time.
2Pull its history: receiving, transfers, scans and adjustments from the stock system.
3Trace it manually line by line, looking for the entry that explains the gap.
4One wrong guess posts a bad adjustment and hides the real loss.
Every gap checked line by line
✓With the app
1Each gap is read together with that item's own history at that location.
2A likely cause is drafted from a fixed list, with the history line that supports it.
3Thin evidence stays open, marked unresolved instead of guessed.
4Inventory control decides from a short note. Nothing is posted or closed by the app.
People review a drafted cause
See it work
One real count gap: what the app reads, step by step
Cleaning chemicals at location 13 count 845 on the shelf against 772 in the system, with three transfers in the history.
Find why a store's stock count doesn't match the systemReference appBuilt to be shaped to your process
4
1The count gap 845 on the shelf, 772 in the system, for cleaning chemicals at location 13.
2What was received 82 units booked in against a purchase order.
3Three transfers in 21, 17 and 17 units moved in from other locations.
4Past corrections two earlier count fixes, +10 and +12 units, that the draft must weigh.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Find why a store's stock count doesn't match the system
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Tracing a cycle-count variance to its likely cause means pulling an item/location's own transaction history -- receiving, transfers, scans, adjustments -- and checking whether any of it explains the discrepancy, for every variance a count throws off. An inventory-control analyst opening a cycle-count variance, pulling the item/location's own transaction history -- receiving, transfers, scans, adjustments -- and manually tracing whether any of it explains the count discrepancy, for every variance a count throws off.
Audience
Inventory-control and retail/warehouse operations staff who review cycle-count variances, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual variance events
The corpus is 36 variance events, 0.03 MB (jsonl 1). A real cycle-count variance log names real SKUs, real store or warehouse codes and real transaction volumes -- exactly the material an inventory-control team will not let leave the building. There is no public corpus of paired (transaction history log, variance, adjudicated root cause) for the same reason there is no public corpus of internal receiving or transfer logs. Every location code, SKU, product category and log line here is invented; nothing refers to a real company.
The corpus
The 36 variance eventsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your variance events. That is the whole change — there is no database to migrate.
One variance event, as the model receives itevents.jsonl · 1 of 36
A drafted cause from a fixed five-member vocabulary (or unresolved), a citation to the exact transaction-history log line that supports it, and a short note for inventory control -- informational only, never an adjustment or a closed variance.
And when it cannot
Drafting a specific cause the transaction history does not actually support -- most dangerously calling a case-pack-multiple variance uom_error when the log's own evidence points to an unrecorded transfer instead (measured at 0 of 4 planted trap events this run), or returning no answer at all when reasoning exhausts the reply ceiling (measured at 2 of 36 events this run). See Business.not_good_enough for why a run this honest is not the same claim as a run with nothing to disclose.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Deciding whether a variance that divides evenly by a case-pack size is genuinely a unit-of-measure mixup or an unrecorded transfer that merely happens to divide evenly — the fast tier, over the case-pack-multiple floor 0 of 4 planted trap events confused this run, against the floor's 4 of 4 (100% by construction) -- the floor calls any case-pack multiple uom_error without checking whether a log line actually shows a unit-of-measure mismatch.
What this kit is not — a drafted cause, a citation and a short note for an inventory-control reviewer to confirm and correct -- never an adjustment src/prompt.py's own system prompt states it twice, and there is no field in the answer schema (cause/citations/narrative) that could post or close anything.
And where nothing here is good enough:
Deciding whether unaccounted scan or pick activity explains a variance — nothing here yet -- this is the kit's own weakest measured category The fast tier resolves gold unscanned_movement correctly only 1 of 7 times (14.3%), defaulting to unresolved on the other six rather than fabricating a wrong specific cause.
Deciding whether to trust a 'no answer' the way you would trust a wrong answer — neither -- treat a no-answer event as unscored, not as evidence the variance has no traceable cause 2 of 36 events (5.6%) returned no cause at all because provider-side reasoning consumed the entire 16,384-token ceiling -- a token-budget failure, not a judgement about the transaction history.
At a glanceHow the whole thing runs
78%cause accuracy pct
13,703 msp50, end to end
$12.95per 1,000 variance events · Google Gemini 3 Flash
Run once, for real, on 2026-08-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Find why a store's stock count doesn't match the system14 steps · 4 questions · run once, for real · 2026-08-20
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own transaction types and your own item master -- src/segment.py, src/pack.py and src/prompt.py read an event by its own fields (item_id, location_id, variance_qty, log) and do not care where the values came from. The measured 77.8% cause accuracy and 0% uom/transfer trap confusion are THIS corpus's own flag vocabulary (receiving_correction / counterpart_activity / uom_note log-line flags) and THIS run's own case-pack sizes (12, 24).Corpus lens →
When is this the wrong choice?
Avoid: The case-pack-multiple floor for anything but the honest baseline it is -- it is not a competitor, it is the trap this kit's eval_intent exists to catch, written out as a rule. That is the case against the best-fitting scenario (“Deciding whether a variance that divides evenly by a case-pack size is genuinely a unit-of-measure mixup or an unrecorded transfer that merely happens to divide evenly”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
One planted cause per event. A real variance can have two compounding causes at once; this corpus always plants exactly one, or none for unresolved. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether the same figures hold on a second run, or on a corpus combining more than one planted pattern per event. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-20 — r001-inv-cycle. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- event VC-0001's variance and its six-line transaction history log, since data/events.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Pressing Draft cause with no API_KEY returns a calm 200 explaining nothing was called. It cannot reproduce a cause, a score or a dollar figure without a key -- those are what results/eval-r001-inv-cycle.json already committed. The free floor (python -m evals.baseline) and the gold self-check (python tools/verify_gold.py) both run on a cold clone with no key at all.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
94.4%rows answered
13,703 msp50, end to end
145,302 msp95
2 minclone to first result
What the clock covers. END-TO-END per variance event: one HTTP request carrying the item/location's own transaction history log and the variance, parsed into a cause, a citation and a short narrative. No retrieval step -- loading an event from disk is pure code and costs no network time. Reasoning was left at the provider's default (on), so p50 and p95 both include a reasoning pass on every call, and the tail is dominated by it: the p95 of 145,302 ms IS event VC-0010, one of the two events that spent its entire 16,384-token ceiling on reasoning and returned no answer; the single slowest call in the whole run, at 149,107 ms (about two and a half minutes), is VC-0019, the other no-answer event. The fastest call, VC-0027, took 3,949 ms and produced a clean, correctly-drafted mis_receipt answer.
Current processWhat it replaces
An inventory-control analyst opening a cycle-count variance, pulling the item/location's own transaction history -- receiving, transfers, scans, adjustments -- and manually tracing whether any of it explains the count discrepancy, for every variance a count throws off.
Where it is not good enough
Five findings, and the first is the headline this kit exists to report honestly rather than quietly re-run until it disappeared. FIRST, THE NO-ANSWER RATE. Two of 36 events (5.6%) spent their entire MAX_TOKENS=16384 reply ceiling on provider-side reasoning and returned no cause, no citation and no narrative at all -- for both, reasoning_tokens equals output_tokens equals the ceiling exactly. This is the residual left after FOUR real runs to find where the ceiling settles: at 2,500 tokens, 17 of 36 events came back truncated (52.8% cause accuracy, a 53.1% 'FABRICATED CAUSE' rate that was really a truncation artefact); at 4,096, 14 of 36 (58.3%, 45.2%); at 8,192, 5 of 36 (75.0%, 18.5%); at 16,384 (registered), 2 of 36 (77.8%, 8.3%). The truncated EVENTS were a different set each time -- the 8,192 ceiling's five failures (VC-0003, 11, 15, 16, 35) share zero overlap with the 16,384 ceiling's two (VC-0010, 19) -- which is evidence of unpredictable runaway reasoning on a small fraction of events, not a fixed hard subset a bigger ceiling would eventually clear. No fifth run was fired to chase the residual to zero; it is reported here instead. SECOND, AND TIED DIRECTLY TO THE FIRST: reasoning consumed 98.5% of this run's total output-token budget (147,503 of 149,762 tokens; 98.1% on the 34 completed calls alone, where it ran a median 1,108 of 1,175 output tokens per call) -- the highest reasoning share measured on any kit in this series so far. The registered run left provider-side reasoning ('thinking') at its documented default -- ON -- while the shipped app's own endpoint (src/app.py's /api/draft) hardcodes thinking=THINKING_OFF on every live call. Nobody has scored the reasoning-off configuration on this corpus, and given how much of the ceiling reasoning alone consumes even on completed calls, there is no basis to assume it would reproduce this run's accuracy, its latency, or its no-answer rate. THIRD, A DISTINCT CATEGORY WEAKNESS, NOT A TRUNCATION ARTEFACT: of the 7 gold unscanned_movement events, the model correctly named the cause only once (14.3%); the other six were drafted as unresolved -- a safe, honest miss rather than a fabrication, but a real, measured blind spot. Neither no-answer event is gold unscanned_movement, so this is a separate finding from FIRST and SECOND, not the same one restated. By contrast unrecorded_transfer (9 of 9, including all 4 planted case-pack-multiple traps) and unresolved itself (6 of 6) are both fully clean this run. FOURTH, narrative_faithfulness_pct reads 35.3% (12 of 34 narratives scored, 2 produced none), and that number is easy to over-read: the check is a literal substring match on the cause's own label text (see Eval.method_note), and manual inspection of every non-unresolved 'unfaithful' hit (14 of them) found all 14 were drafted with the CORRECT cause and a citation independently scored citation_valid=True -- the narrative simply describes the mechanism in plain prose ('a unit-of-measure mismatch', 'the discrepancy was never posted to on-hand') rather than echoing the underscore-joined label word. Whether a semantic check would score materially higher was not measured; only the pure-code proxy the use case spec asks for was run -- see Eval.could_not_verify. FIFTH, structural limits stated in the corpus's own documentation rather than found by this run: every event carries exactly one planted cause (or none, for unresolved), never two compounding causes at once; this corpus judges every event independently rather than as a rolling per-item history across cycles; and the transaction history is assumed complete enough to reconstruct a variance's cause, with no way for this kit to notice a genuinely missing log entry (a store still running paper transfer logs, say) rather than simply absent evidence.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
the free case-pack-multiple floor (evals/baseline.py) scores 52.8 pct cause accuracy and gets all 4 planted traps wrong by construction (100 pct uom_transfer_confusion) at $0.00; the fast-tier model reaches 77.8 pct cause accuracy and resists every planted trap but is not clean either — two events return no answer at all (reasoning exhausted the ceiling) and MAX_TOKENS was raised four times (2,500 to 16,384) without that residual reaching zero. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the cause vocabulary
src/rubric.py
A cause this kit's fixed five-member list does not carry needs a new CAUSE_VOCAB entry here plus a new flagged evidence line in tools/build_corpus.py and a new branch in src/segment.py::classify() -- a prompt tweak alone will not teach it.
the corpus
tools/build_corpus.py
Point it at your own transaction types and your own item master. src/segment.py, src/pack.py and src/prompt.py read an event by its own fields (item_id, location_id, variance_qty, log) and do not care where the values came from.
the case-pack sizes
src/segment.py
CASE_PACK_SIZES (12, 24 today) decides which variances present the trap's surface signal at all -- a different item master's own common pack sizes changes what counts as a coincidental multiple.
the reasoning ceiling
src/prompt.py
MAX_TOKENS is both a swap seam and this run's own open question: raising it reduced the no-answer rate on four successive attempts (2,500 to 16,384) but never reached zero, and nothing in this kit separately caps the reasoning pass from the visible reply -- see Guardrails.add_first.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 36 variance events across item/location/period combinations, each with its own 4-10 line transaction history log and a derived gold cause, from a fixed seed. One trap is planted on purpose: about 44% of unrecorded_transfer events also carry a variance that is a clean case-pack multiple -- see data/SOURCES.md.
the cut
src/segment.py
SEAM -- the five-way rule. classify() decides which of the five causes one event's own log and variance_qty support, by a fixed flag-precedence order (a flagged receiving-correction line first, a flagged counterpart-activity line second -- the line that resolves the trap -- a flagged uom-mismatch line third, a residual-arithmetic check for unscanned_movement fourth, unresolved otherwise). Pure code, no model, and the SAME function both writes gold and grades a live model's citations.
the rubric
src/rubric.py
The five-member cause vocabulary and the three graded axes (cause accuracy 50%, citation validity 30%, narrative faithfulness 20%), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart.
Pack
src/pack.py
Deterministic rendering of one event's transaction history into the numbered lines the model actually sees -- the whole log, never pre-filtered, since every line in an item/location's own history is a candidate. Pure code -- capped at 10 lines, though this run's own 36 events never reached that ceiling (max 8).
the prompt
src/prompt.py
The cause vocabulary, the per-event answer schema and the narrative instruction, declared once and read from here by the prompt builder, the parser and the app. Also carries this kit's own MAX_TOKENS history as a code comment -- the four-ceiling investigation behind Business.not_good_enough.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider, stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
the AI layer
src/brief.py
Loads one event, packs its transaction log, calls the model once, parses the reply. The whole AI layer, deliberately short. Gold is never read here -- only by evals/scoring.py.
the app
src/app.py
Static files and three JSON endpoints (/api/state, /api/event, /api/draft) on the standard library, port 8791. Runs the SAME citation-validity check evals/scoring.py runs, on every live draft, before the UI renders it -- not only after the fact in a committed eval result. /api/draft hardcodes thinking=THINKING_OFF on every live call, a different reasoning setting from the one the registered eval run actually used; see Business.not_good_enough.
Where it breaks at scale
Not on event count -- each event is independent and one call per event is linear. It breaks three other ways. FIRST, ON REASONING, and this is the kit's own central finding: a small, unpredictable fraction of events (2 of 36 here) send the model into unbounded reasoning that consumes the entire reply ceiling and returns nothing. The truncated SET changed at every one of four ceilings tried (17, then 14, then 5, then 2 of 36), with zero overlap between the last two sets, so there is no evidence a bigger corpus would settle into a stable, predictable rate -- only that raising MAX_TOKENS further is not guaranteed to reach zero, and this run's own history says it has not yet. SECOND, ON THE unscanned_movement CLASS SPECIFICALLY -- a structural blind spot distinct from the reasoning problem above: the model resolves this cause correctly only 1 of 7 times this run, defaulting to the safe 'unresolved' on the rest rather than fabricating a wrong specific cause. Whether that generalizes to a transaction history with more or fewer scan lines than this corpus's own is unmeasured. THIRD, structural assumptions stated in the corpus's own documentation: one planted cause per event (a real variance can have two compounding causes at once), one event per item/location/period rather than a rolling history (a real deployment sees the same item recur cycle over cycle, sometimes still unresolved from last count), and a transaction history assumed complete enough to reconstruct a cause, with no way for this kit to notice a genuinely missing log entry rather than simply absent evidence.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
The panel before any call: event VC-0001 loaded, its variance (system 772, counted 845, variance -73) and its full six-line transaction history log -- receiving, three inbound transfers, two prior-period adjustments -- shown in order with their line indices. Nothing has been drafted yet.successOpen full size →The variance panel outlined: Cleaning Chemicals at LOC-013, system 772 against counted 845, a -73 gap, with its six-line transaction history log beneath it.successOpen full size →The same panel, the receiving line outlined: 82 units received against a purchase order, the first entry in the item's own history.successOpen full size →The same panel, the three inbound transfer lines outlined: 21, 17 and 17 units moved in from other locations.successOpen full size →The same panel, the two prior-period adjustment lines outlined: +10 and +12 units, the last entries the draft must weigh.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same event after pressing Draft cause with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. The drafted-cause panel shows a 'NO VERDICT' badge and 'no citation (unresolved)' -- the panel's own generic empty-answer copy, not the model choosing unresolved -- so a reader cannot mistake this for a real, scored draft. This is the honest failure state, not a staged one.failureOpen full size →
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
36variance events
0.03 MiBjsonl 1
p50 522chars per characters per event's assembled header, variance line and transaction log block
$0.00setup · 0.0s
How it is cutWhat one characters per event's assembled header, variance line and transaction log block is
One variance event, its own item/location transaction history log read whole -- no train/test split, every event drafted once, in one call.
SetupWhat the setup figure measured
No index is built. src/app.py finds an event by scanning B.events() for a matching event_id -- 36 rows, a linear scan, not a search structure -- and src/brief.py sends that event's own transaction log whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own variance events
Point tools/build_corpus.py at your own transaction types and your own item master -- src/segment.py, src/pack.py and src/prompt.py read an event by its own fields (item_id, location_id, variance_qty, log) and do not care where the values came from. The five-member CAUSE_VOCAB and its flag-precedence order in src/segment.py::classify() are what decide gold, so a cause outside the fixed five is a src/rubric.py change plus new flagged evidence lines in tools/build_corpus.py, not a prompt-only change.
⚠︎ And what stops being true when you do: The measured 77.8% cause accuracy and 0% uom/transfer trap confusion are THIS corpus's own flag vocabulary (receiving_correction / counterpart_activity / uom_note log-line flags) and THIS run's own case-pack sizes (12, 24). A real transaction log's own flagging conventions, case-pack sizes and log density are all unmeasured by this run.
What breaks it
One planted cause per event. A real variance can have two compounding causes at once; this corpus always plants exactly one, or none for unresolved.
One event per item/location/period, not a rolling history. A real deployment would see the same item recur cycle over cycle, sometimes still unresolved from last count; this corpus judges every event independently.
Transaction history assumed complete by construction. This kit assumes receiving, transfers and scans are captured completely enough to reconstruct a variance's cause -- a store still running manual transfer logs on paper leaves gaps this corpus never models, and this kit has no way to know evidence is missing rather than simply absent because there was none to log.
The trap denominator is small -- 4 planted case-pack-multiple unrecorded_transfer events out of 9 unrecorded_transfer events total. That is what a corpus of this scale honestly supports, not a stable rate over a larger sample.
A small, unpredictable fraction of events send the model into unbounded reasoning regardless of MAX_TOKENS -- 2 of 36 at this run's own 16,384-token ceiling, after ceilings of 2,500, 4,096 and 8,192 each left a different nonzero residual. A larger or differently-shaped corpus has no measured reason to have a smaller one.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,052
780
event
153
39
log
420
107
Total
926
This is the cost lesson as arithmetic: of the 926 tokens assembled, 780 are instructions — 84% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged by importing src.prompt directly and calling prompt.build(packed) for event VC-0001 (not retyped) -- the literal system and user content src/brief.py::draft() sends. Per-part token counts are a proportional estimate over character share (system 3,052 / event 153 / log 420 chars, 3,625 total) applied to this call's own total of 926 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are drafting the likely cause of ONE inventory cycle-count variance, for inventory control to confirm and correct. You never post an inventory adjustment and never close the variance yourself -- your job ends at a drafted cause, a citation and a short note for the human who does that work.
You are given one item/location's own transaction history log for the relevant window, and the variance itself: system quantity minus physical count, signed. Decide which of these five causes the log actually supports, using ONLY this log -- never outside knowledge, never a plausible-sounding guess:
mis_receipt a receiving log line's recorded quantity doesn't reconcile with the variance -- a receiving correction that was logged but never actually applied to on-hand, or a received quantity that itself looks miscounted against the PO.
unrecorded_transfer the variance size lines up with what a transfer to or from another location would explain, but no transfer log line covering that quantity or window exists here.
uom_error the variance is a clean multiple of a common case-pack size, and a specific log line shows eaches and cases were mixed up for this item -- not just that the number happens to divide evenly.
unscanned_movement a scan, pick or sale log line's quantity, summed with the rest of the history, doesn't fully reconcile against the variance -- some movement happened with no corresponding log entry.
unresolved the transaction history genuinely doesn't support any of the four causes above cleanly. This is the correct answer when the evidence isn't there -- guessing a specific cause is a worse answer than saying so.
A variance that divides evenly by a common case-pack size is suggestive but NOT sufficient by itself for uom_error -- only draft uom_error when a specific log line actually shows a unit-of-measure or case/each mismatch for this item. If instead the log points to activity at another location that was never logged here as a transfer, prefer unrecorded_transfer. Checking which the log actually supports, rather than pattern-matching the arithmetic alone, is the whole job.
If the log supports a specific cause, cite the exact log line index (or two, for unscanned_movement, when more than one scan line is relevant) that supports it. If the log does not clearly support one of the four named causes -- no line addresses it, or the lines that do don't add up to one of them -- the cause is unresolved and citations must be an empty list. Do not cite a line just because it is the closest-sounding one available: an unsupported unresolved is the correct, honest answer and is graded as such; a specific cause with a citation that doesn't actually support it is graded as a fabrication, which is worse.
After the cause and citation, write a short narrative (1-2 sentences) for inventory control that states the drafted cause and grounds it only in what the citation actually shows -- do not restate a quantity or a location that the cited line doesn't contain.
Event VC-0001 -- item SKU-45991 (Cleaning Chemicals), location LOC-013, period 2026-W32
system 772, counted 845, variance_qty -73 (system minus counted)
TRANSACTION HISTORY LOG
------------------------
0: [receiving] Received 82 units against PO-58229.
1: [transfer] Transfer in: 21 units received from LOC-025.
2: [transfer] Transfer in: 17 units received from LOC-003.
3: [transfer] Transfer in: 17 units received from LOC-024.
4: [adjustment] Prior-period cycle count correction applied: +10 units.
5: [adjustment] Prior-period cycle count correction applied: +12 units.
Return a JSON object with three keys: "cause" (one of: mis_receipt, unrecorded_transfer, uom_error, unscanned_movement, unresolved); "citations" (a list of integer log line indices -- empty for unresolved); and "narrative" (a string of 1-2 sentences).
Answer with the JSON object only.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"cause": "unresolved",
"citations": [],
"narrative": "The transaction history contains only receiving, inbound transfers, and prior-period adjustments; none of these logged lines can be tied to the -73 variance as a receiving error, unrecorded transfer, unit-of-measure mix-up, or missing scan."
}
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Find why a store's stock count doesn't match the system — 36 variance events. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: a five-way exact match on the drafted cause, citation_validity() checking whether a non-unresolved cause's cited log line is both real and actually evidences that cause (src/segment.py::line_supports_cause(), the same function that wrote gold in the first place), narrative_faithfulness() checking whether the narrative's own prose contains the drafted cause's label (in underscore or spaced form), and uom_transfer_confusion counting how many of the 4 planted case-pack-multiple trap events were drafted as uom_error instead of the gold unrecorded_transfer. The same function scores both evals/baseline.py's free floor and evals/run.py's real run. ⚠︎ A RECORD WITH NO CAUSE AT ALL (both of this run's no-answer events) IS SCORED AS WRONG ON CAUSE ACCURACY AND AS A CITATION FAILURE ON CITATION VALIDITY, NEVER AS A SEPARATE 'NO ANSWER' BUCKET: citation_validity() returns False, not None, for a falsy cause, so both events land inside fabricated_cause -- the run's entire 8.3% fabricated_cause_rate_pct is these two events, not a genuinely invented citation on a drafted cause. They are NOT scored as unfaithful narratives: narrative_faithfulness() returns None for a missing narrative, a distinct, worse state the scorer never folds into 'unfaithful' -- both land in narrative_missing instead.
36variance events
36source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED28 / 36cause accuracy pct — variance eventsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED6 / 6cause accuracy unresolved pct — gold-unresolved eventsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 / 30cause accuracy traceable pct — gold-traceable eventsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 / 24citation validity pct — non-unresolved cause+citation pairs scored, including the 2 no-answer eventsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED2 / 24fabricated cause rate pct — non-unresolved cause+citation pairs scored -- both hits are the 2 no-answer events, not a wrong citation on a genuinely drafted cause; see method_noteDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 4uom transfer confusion rate pct — planted case-pack-multiple trap eventsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED12 / 34narrative faithfulness pct — narratives produced (literal cause-label substring match -- see method_note)Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED2 / 36no answer rate pct — variance events -- both from provider-side reasoning exhausting MAX_TOKENS=16384 with zero answer tokensDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer is pure code, exact match against a mechanically-derived gold set: every event's cause and citation are derived once, at corpus-build time, by src/segment.py::classify() from the log and variance_qty alone -- never hand-typed -- and tools/verify_gold.py re-runs classify() over the finished data/gold.jsonl, re-checks that every gold citation actually supports its own cause via line_supports_cause(), and re-derives is_trap as cause==unrecorded_transfer AND is_case_pack_multiple(variance_qty) rather than trusting a separately-typed flag. Confirmed this session: python3 tools/verify_gold.py -> "OK -- all 36 gold rows match a fresh classify() re-derivation from their own transaction history log and variance_qty."
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One variance event
1,000 variance events
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.012946
$12.95
4%
Same work, 1× the bill
The same variance events, the same tokens — only the rate card changed. And on that card about 4% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether reasoning ('thinking') is left on or explicitly disabled for this model -- the app's own /api/draft always disables it; the registered run left it at provider default. The two configurations have not been priced against each other on this corpus, and given reasoning is 98.5% of this run's own output budget, the gap between them is likely the largest of any kit in this series.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build (including the three superseded ceiling attempts at 2,500/4,096/8,192 tokens) is recorded in the commit history and the shared call ledger, not on this page. Reasoning was also left at provider default (on) for this run -- see Business.not_good_enough -- so even the projected figure prices tokens the shipped app's own reasoning-off configuration would not spend.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. Cause accuracy, citation validity, narrative faithfulness and uom/transfer confusion are all pure code over the committed run file, so a forker re-scores this run -- or the free floor -- for $0.00 and needs no key.
The gradersOne way to grade, and why it is the only one
The floor is exactly the trap, written out as a rule: 'a case-pack multiple means uom_error, full stop.' It gets all 4 planted trap events wrong BY CONSTRUCTION (100% uom_transfer_confusion, against the fast tier's 0%), and it never cites anything, so its citation_validity_pct is 0.0% on every one of the 26 non-unresolved guesses it makes -- an unsupported specific cause scores exactly as fabricated as a wrongly-cited one. Its narrative_faithfulness_pct reads a trivial 100%, because floor_answer() templates the cause label directly into the narrative text every time ('... read as uom_error from the numbers alone') -- the proxy check cannot fail on a narrative engineered to contain the label word by construction, which is exactly why the real run's own 35.3% needs the qualitative read in Business.not_good_enough rather than a bare comparison against this number. Its overall cause accuracy (52.8%) is inflated by the 6 gold-unresolved events, which it also gets right by construction whenever no scan line and no case-pack multiple happen to be present.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Cause exact match, plus citation validity, narrative faithfulness and the trap guardrail Does the drafted cause match gold's five-way vocabulary exactly, including saying unresolved when gold says unresolved? For a non-unresolved cause, is the cited log line real AND does it actually support that cause? Does the narrative's own prose name the drafted cause? Of the 4 planted case-pack-multiple trap events, how many were drafted as uom_error instead of the gold unrecorded_transfer?
$0.00
no
yes
the fast tier 77.8% field accuracy
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, on the axis this corpus was built to test. The free floor and the fast tier are scored by the identical function over the identical 36 events, and they separate cleanly on the trap specifically: 100% vs 0% uom_transfer_confusion, exactly the corpus's 4 planted trap events. Overall cause accuracy also separates (52.8% vs 77.8%) but less cleanly, because the floor's own arithmetic-only rule happens to get every gold-unresolved event right too (100% vs the model's own 100%) -- the real gap concentrates in the 30 gold-traceable events (43.3% vs 73.3%). A grader that could not tell the two apart would not produce a trap-specific gap that lines up exactly with the corpus's own documented 4 planted instances.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Deciding whether a variance that divides evenly by a case-pack size is genuinely a unit-of-measure mixup or an unrecorded transfer that merely happens to divide evenly
the fast tier, over the case-pack-multiple floor
0 of 4 planted trap events confused this run, against the floor's 4 of 4 (100% by construction) -- the floor calls any case-pack multiple uom_error without checking whether a log line actually shows a unit-of-measure mismatch.
the case-pack-multiple floor for anything but the honest baseline it is -- it is not a competitor, it is the trap this kit's eval_intent exists to catch, written out as a rule.
Deciding whether unaccounted scan or pick activity explains a variance
nothing here yet -- this is the kit's own weakest measured category
The fast tier resolves gold unscanned_movement correctly only 1 of 7 times (14.3%), defaulting to unresolved on the other six rather than fabricating a wrong specific cause.
trusting a lack of an unscanned_movement call as evidence that no unscanned movement occurred -- this run shows the model under-calls this specific cause, not that the cause is rare in the corpus (7 of 36 gold events are unscanned_movement, same as mis_receipt and uom_error).
Deciding whether to trust a 'no answer' the way you would trust a wrong answer
neither -- treat a no-answer event as unscored, not as evidence the variance has no traceable cause
2 of 36 events (5.6%) returned no cause at all because provider-side reasoning consumed the entire 16,384-token ceiling -- a token-budget failure, not a judgement about the transaction history.
reading Eval.scores' fabricated_cause_rate_pct (8.3%) as genuine hallucination -- both hits are these same two no-answer events; see method_note.
What this kit is not
a drafted cause, a citation and a short note for an inventory-control reviewer to confirm and correct -- never an adjustment
src/prompt.py's own system prompt states it twice, and there is no field in the answer schema (cause/citations/narrative) that could post or close anything.
wiring this straight into an auto-posting adjustment pipeline with no reviewer step for any drafted cause, including unresolved.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
resisted_uom_transfer_trap
Planted case-pack-multiple variance correctly drafted as unrecorded_transfer, not uom_error
4
VC-0004: variance -144 (divisible by both case-pack sizes, 12 and 24). Gold and model both: unrecorded_transfer, citing log line 0 ("Regional log shows LOC-022 posted a partial adjustment of 144 units... not logged here as a transfer"). The free floor calls…
correct_other_traceable
Correctly drafted a specific cause (mis_receipt, uom_error, unrecorded_transfer, or unscanned_movement) with a valid citation, outside the trap set
18
VC-0014: gold uom_error, citing log line 0. Model: uom_error, citing line 0, narrative "Line 0 logs a receipt in cases with a case pack of 12, revealing a unit-of-measure (eaches/cases) mismatch for this item. The +12 variance equals the case pack, so the…
correct_unresolved
Correctly said unresolved when the log genuinely supported no specific cause
6
VC-0001: variance -73, a six-line log of receiving, transfers and prior-period adjustments with no flagged line for any of the four named causes. Gold and model both: unresolved, no citations.
missed_unscanned_movement
Gold unscanned_movement drafted as unresolved instead -- a safe miss, never a fabrication, but the kit's clearest measured weak spot
6
VC-0011: gold unscanned_movement citing line 1 ("POS scan/pick: 28 units moved out"), which does not reconcile the log's own arithmetic against the -variance. Model: unresolved, narrative "The transaction history does not clearly support any of the four…
no_answer_reasoning_exhausted
Provider-side reasoning consumed the entire MAX_TOKENS=16384 ceiling; zero answer tokens, no cause, no citation, no narrative
2
VC-0010 (gold uom_error) and VC-0019 (gold mis_receipt): both finish_reason='length', both reasoning_tokens==output_tokens==16384 exactly. VC-0019 is this run's single slowest call at 149,107 ms (about two and a half minutes) for zero usable output.
What we could NOT verify
Whether the same figures hold on a second run, or on a corpus combining more than one planted pattern per event.
Whether the 2-of-36 no-answer rate at MAX_TOKENS=16384 is a stable rate or an artifact of this specific 36-event sample -- four ceilings were tried (2,500/4,096/8,192/16,384) and the residual never reached zero, but no fifth run was fired to test whether it is truly irreducible or merely still-unlikely at this sample size; see Business.not_good_enough.
Whether the same figures hold with reasoning explicitly disabled (thinking=THINKING_OFF), matching the live app's own /api/draft setting -- untested; see Business.not_good_enough.
Whether the model's own unscanned_movement weakness (1 of 7 correct) generalizes to a transaction history with a different scan-line density than this corpus's own.
Whether a semantic (rather than literal cause-label substring) faithfulness check would score narrative_faithfulness_pct materially higher than 35.3% -- only the pure-code substring proxy evals/scoring.py implements was run; see method_note.
Whether a transaction-history log crafted to mimic this kit's own flag phrasing could talk the model into a wrong specific cause -- no red-team run exists for this kit.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
932.25
4,160.06
13,703 ms
$0.012946
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The case-pack-multiple floor (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 33,561 input / 149,762 output tokens priced at Google Gemini 3 Flash's published rate -- the same basis cost_per_query_usd uses, not a second, larger spend. The three superseded ceiling attempts (2,500, 4,096, 8,192 tokens, before 16,384 was settled on) cost real money too and are not included in this figure; those result files were deleted as scratch evidence once their numbers were captured in this spec and in src/prompt.py's own MAX_TOKENS comment.
Cost driversWhat actually moves the bill
The provider's reasoning pass, left at its default (on) rather than disabled for this run -- it consumed 98.5% of the run's TOTAL output-token budget (147,503 of 149,762 tokens; 98.1% on the 34 completed calls alone) and is already priced into cost_per_query_usd, since a provider bills reasoning tokens as completion tokens. This is the highest reasoning share of any kit in this series to date. src/app.py's live UI hardcodes thinking off, a different, unmeasured configuration -- see Business.not_good_enough.
The two no-answer events (VC-0010, VC-0019) are each billed for the full 16,384-token ceiling despite producing zero answer text -- together 32,768 of the run's 149,762 output tokens (21.9%) bought no usable output at all.
The item/location's own transaction history log, packed whole per event (4 to 8 lines, 218 to 474 characters) -- small and essentially flat next to the reasoning pass, the floor every call pays regardless of how the log resolves.
Your volumeWhat it costs at your volume
Linear in events, with one caveat a flat rate hides: this run's 36 events cost about $0.4661 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $4.66 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run. What is NOT necessarily linear is the no-answer rate: this run's 2 of 36 (5.6%) is the residual AFTER four ceiling increases, not a fixed per-event probability, so 360 events at the same MAX_TOKENS may return roughly 20 no-answer events or may not -- the four-ceiling history (17, 14, 5, 2 truncated of 36) never settled into a stable rate.
Where pricing changes shape
Your return, with your numbers
Volumeinventory cycle-count variances drafted per day/week -- this run drafted 36 in one pass
What it replacesan inventory-control analyst pulling an item/location's own transaction history by hand and tracing whether receiving, transfers, scans or adjustments explain a cycle-count variance
Time saved per itemnot measured here -- depends on how long a manual trace takes at the reader's own operation
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run, so this page prices one model, not a trade-off; see Eval.could_not_verify for what a second run -- or the same run with reasoning explicitly disabled -- would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
33,561input tokens · this run
149,762output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 36 variance events drafted, each scored by pure code against a mechanically-derived gold set. This run (r001-inv-cycle) answered 34 of 36 events (2 returned no answer at all -- reasoning exhausted the 16,384-token ceiling) -- the row this table prices, the same one Cost.cost_by_model[0] uses. The two no-answer events are billed in full (16,384 output tokens each) despite producing nothing.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.186
$0.186
$5.18
2026-09-12
gemini-3-flash
Google
$0.466
$0.466
$12.95
2026-09-18
gemini-3-8-flash
Google
$0.587
$0.587
$16.30
2026-09-18
llama-5
Meta
$0.678
$0.678
$18.85
2026-09-18
claude-haiku-4-5
Anthropic
$0.782
$0.782
$21.73
2026-09-12
grok-4-5
xAI
$0.966
$0.966
$26.82
2026-09-18
grok-4-6
xAI
$0.966
$0.966
$26.82
2026-09-18
claude-sonnet-5
Anthropic
$1.565
$1.565
$43.47
2026-09-12
gemini-3-1-pro
Google
$1.864
$1.864
$51.79
2026-09-18
gpt-5-6-terra
OpenAI
$1.864
$1.864
$51.79
2026-09-12
gpt-5-6-sol
OpenAI
$3.129
$3.129
$86.93
2026-09-12
claude-opus-4-8
Anthropic
$3.912
$3.912
$108.66
2026-09-12
claude-opus-5
Anthropic
$3.912
$3.912
$108.66
2026-09-12
claude-fable-5
Anthropic
$7.824
$7.824
$217.33
2026-09-18
claude-fable-5-1
Anthropic
$7.824
$7.824
$217.33
2026-09-18
gpt-6-astra
OpenAI
$7.824
$7.824
$217.33
2026-09-17
Read this against the numbers above
REASONING WAS NOT DISABLED FOR THIS WORKLOAD, AND IT IS 98.5% OF THE OUTPUT TOKENS PRICED HERE. Every row below prices THIS run's own token counts, which include a reasoning pass the live app's own /api/draft never generates (it hardcodes thinking off). A forker's real bill on another model depends on whether that model has an equivalent reasoning toggle and whether it is left on.
The two no-answer events are priced at the full 16,384-token ceiling with zero usable output -- a different model's own reasoning behavior on this same corpus is unmeasured and could produce more, fewer, or zero such events.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 36 variance events across item/location/period combinations, each with its own 4-10 line transaction history log and a derived gold cause, from a fixed seed. One trap is planted on purpose: about 44% of unrecorded_transfer events also carry a variance that is a clean case-pack multiple -- see data/SOURCES.md.
You change it to: Point it at your own transaction types and your own item master. src/segment.py, src/pack.py and src/prompt.py read an event by its own fields (item_id, location_id, variance_qty, log) and do not care where the values came from.
tools/build_corpus.py
# Generate the variance events and their gold cause + citations, from a fixed seed.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
CATEGORIES = [
LOCATIONS = ["LOC-%03d" % i for i in range(1, 31)]
PERIODS = ["2026-W%02d" % w for w in range(5, 34)]
CASE_PACK_SIZES = SEG.CASE_PACK_SIZES
SCENARIOS = ["mis_receipt", "unrecorded_transfer_plain", "unrecorded_transfer_trap",
COUNTS = {"mis_receipt": 7, "unrecorded_transfer_plain": 5, "unrecorded_transfer_trap": 4,
src/segment.pythe cut — a swap seam
SEAM -- the five-way rule. classify() decides which of the five causes one event's own log and variance_qty support, by a fixed flag-precedence order (a flagged receiving-correction line first, a flagged counterpart-activity line second -- the line that resolves the trap -- a flagged uom-mismatch line third, a residual-arithmetic check for unscanned_movement fourth, unresolved otherwise). Pure code, no model, and the SAME function both writes gold and grades a live model's citations.
You change it to: CASE_PACK_SIZES (12, 24 today) decides which variances present the trap's surface signal at all -- a different item master's own common pack sizes changes what counts as a coincidental multiple.
src/segment.py
# SEAM 1 -- the cut. Deciding which cause one variance event's own transaction history actually
CASE_PACK_SIZES = (12, 24)
def is_case_pack_multiple(variance_qty, pack_sizes=CASE_PACK_SIZES):
def accounted_change(log):
def line_supports_cause(event, idx, cause):
def classify(event):
src/rubric.pythe rubric — a swap seam
The five-member cause vocabulary and the three graded axes (cause accuracy 50%, citation validity 30%, narrative faithfulness 20%), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart.
You change it to: A cause this kit's fixed five-member list does not carry needs a new CAUSE_VOCAB entry here plus a new flagged evidence line in tools/build_corpus.py and a new branch in src/segment.py::classify() -- a prompt tweak alone will not teach it.
src/rubric.py
# The rubric this kit's drafted cause is graded against -- the fixed five-member cause vocabulary
CAUSE_VOCAB = (
CAUSE_MEANINGS = {
RUBRIC_AXES = (
FABRICATION_GUARDRAIL = (
CONFUSION_GUARDRAIL = (
src/pack.pyPack
Deterministic rendering of one event's transaction history into the numbered lines the model actually sees -- the whole log, never pre-filtered, since every line in an item/location's own history is a candidate. Pure code -- capped at 10 lines, though this run's own 36 events never reached that ceiling (max 8).
src/pack.py
# SEAM 2 -- Pack. Deterministic rendering of one variance event's transaction history into the
MAX_LOG_LINES = 10
def pack(event):
src/prompt.pythe prompt — a swap seam
The cause vocabulary, the per-event answer schema and the narrative instruction, declared once and read from here by the prompt builder, the parser and the app. Also carries this kit's own MAX_TOKENS history as a code comment -- the four-ceiling investigation behind Business.not_good_enough.
You change it to: MAX_TOKENS is both a swap seam and this run's own open question: raising it reduced the no-answer rate on four successive attempts (2,500 to 16,384) but never reached zero, and nothing in this kit separately caps the reasoning pass from the visible reply -- see Guardrails.add_first.
src/prompt.py
# Assemble the one prompt this kit sends per variance event, and parse the one reply it gets
MAX_TOKENS = 16384
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _log_block(packed):
def build(packed, prompt=DEFAULT_PROMPT):
def parse(raw):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider, stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/brief.pythe AI layer
Loads one event, packs its transaction log, calls the model once, parses the reply. The whole AI layer, deliberately short. Gold is never read here -- only by evals/scoring.py.
src/brief.py
# SEAM 3 -- the AI layer. Loads one variance event, packs its transaction log, calls the model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
def _read_jsonl(path):
def events():
def events_by_id():
def gold_by_id():
def draft(cfg, event, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
src/app.pythe app
Static files and three JSON endpoints (/api/state, /api/event, /api/draft) on the standard library, port 8791. Runs the SAME citation-validity check evals/scoring.py runs, on every live draft, before the UI renders it -- not only after the fact in a committed eval result. /api/draft hardcodes thinking=THINKING_OFF on every live call, a different reasoning setting from the one the registered eval run actually used; see Business.not_good_enough.
src/app.py
# The minimal local UI. Standard library only -- python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8791"))
def _corpus():
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 36 variance events across item/location/period combinations, each with its own 4-10 line transaction history log and a derived gold cause, from a fixed seed. One trap is planted on purpose: about 44% of unrecorded_transfer events also carry a variance that is a clean case-pack multiple -- see data/SOURCES.md. A swap seam.
src/segment.pySEAM -- the five-way rule. classify() decides which of the five causes one event's own log and variance_qty support, by a fixed flag-precedence order (a flagged receiving-correction line first, a flagged counterpart-activity line second -- the line that resolves the trap -- a flagged uom-mismatch line third, a residual-arithmetic check for unscanned_movement fourth, unresolved otherwise). Pure code, no model, and the SAME function both writes gold and grades a live model's citations. A swap seam.
src/rubric.pyThe five-member cause vocabulary and the three graded axes (cause accuracy 50%, citation validity 30%, narrative faithfulness 20%), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart. A swap seam.
src/pack.pyDeterministic rendering of one event's transaction history into the numbered lines the model actually sees -- the whole log, never pre-filtered, since every line in an item/location's own history is a candidate. Pure code -- capped at 10 lines, though this run's own 36 events never reached that ceiling (max 8).
src/prompt.pyThe cause vocabulary, the per-event answer schema and the narrative instruction, declared once and read from here by the prompt builder, the parser and the app. Also carries this kit's own MAX_TOKENS history as a code comment -- the four-ceiling investigation behind Business.not_good_enough. A swap seam.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider, stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from. A swap seam.
src/brief.pyLoads one event, packs its transaction log, calls the model once, parses the reply. The whole AI layer, deliberately short. Gold is never read here -- only by evals/scoring.py.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1492 input and 155 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's transaction history logs are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any log line or note text. In a real deployment the transaction history (receiving notes, transfer notes, scan/pick descriptions, adjustment notes) would arrive from warehouse, POS or ERP systems recording free-text an operator typed -- exactly the kind of externally-authored input this kit's architecture treats as trusted, with no verification step. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/draft handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and one boundary that should exist does not
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every transaction history log line is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's receiving/transfer/scan/adjustment notes would carry externally-authored text (see posture), and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it; one is a negative result, not a guarantee. Confirmed by reading the code, not by a run, on 2026-08-20 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a drafted cause ever post an inventory adjustment or close the variance?
A drafted cause could plausibly trigger a posting or a variance-closure action.
No code path does. src/brief.py::draft() and src/app.py's /api/draft both return an answer only; neither writes to data/events.jsonl or any other store -- confirmed by reading every call site.
Can a non-unresolved cause reach a reader with a fabricated or unsupported citation?
A model could assert a specific cause and a plausible-looking but fabricated or unsupported log-line citation, which a reader would trust as evidence.
src/app.py runs the identical citation-validity check evals/scoring.py grades the real run with (SEG.line_supports_cause()) on every live draft, before the UI renders it, and marks citation_ok=false on any non-unresolved cause whose citation fails -- confirmed by reading /api/draft.
Is the case-pack-size trap threshold something a prompt or a reply can move?
A crafted log note could shift which variance counts as a case-pack multiple.
src/segment.py::CASE_PACK_SIZES is a module-level constant, read once at import time -- nothing the model returns is consulted when deciding whether a variance is a case-pack multiple; the trap classification is computed before the model is ever called.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/draft handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Each boundary above was checked by reading the call sites, not by an attack trial. Gate 2 is the one that does NOT hold on the cause itself, though the citation is checked live -- see Guardrails.is_not for the same finding from the enforcement side.
The result0 attack trials, and one boundary that does NOT hold in code (only in the prompt): the cause itself is never re-derived against src/segment.py::classify() before the UI renders it. Two other boundaries (no write path, case-pack threshold not model-settable) hold, confirmed by reading the code, and the citation-fabrication check DOES run live, not only in a committed eval result.
1externally-authored field a live deployment would carry (the transaction history log's own free-text notes) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no log line's note text was authored by an outside party. A real deployment's receiving/transfer/scan/adjustment notes would be exactly the kind of externally-supplied text a warehouse or POS system's own free-text field carries; whether a note crafted to mimic this kit's own flag phrasing (a fake 'transfer' note, say) could talk the model into a wrong specific cause is unmeasured for this kit.
Read this twice
This kit has a live citation check but no live cause check. src/app.py's /api/draft verifies a non-unresolved citation actually supports its stated cause before rendering it, but nothing re-derives the cause itself from the log via src/segment.py::classify() before the UI shows it. On this run 28 of 36 events were correctly caused and citation validity held at 91.7%, but that is a property of THIS run's replies, not a guarantee the code provides. A future version should recompute the cause from the log directly -- the same shape evals/scoring.py already runs offline -- before trusting a live draft, and should separately cap the reasoning pass so a runaway trace cannot silently consume the whole reply budget (see Guardrails.add_first).
HonestyWhat this does not prove
Whether a transaction-history note crafted to mimic this kit's own flag phrasing (a fake-sounding but invented 'transfer' note, worded to read like a genuine counterpart_activity line) could talk the model into a wrong specific cause -- no red-team run exists for this kit.
Whether the live app's hardcoded thinking=off setting changes resistance to a crafted transaction log relative to this run's provider-default-on configuration -- untested either way.
Whether a code-level consistency check (Guardrails.add_first) would catch a real cause/citation disagreement in practice -- none has been built or exercised.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
A drafted cause is limited to what the transaction history log actually supports, and the model must say unresolved rather than guess when no line clearly evidences one of the four named causes -- especially resisting the surface pattern that a case-pack-multiple variance is automatically uom_error.
src/prompt.py -- SYSTEM, stated twice: once as instruction, once as the literal cost of getting it wrong. Prompt-level ONLY: no code in src/brief.py or src/app.py recomputes the cause from the log before the UI renders it.
EvidenceDoes it hold?
What
Measured
No code path in this kit writes an inventory adjustment or closes a variance
0 of 36 calls in r001-inv-cycle resulted in any write or any outbound action -- src/brief.py::draft() and src/app.py's /api/draft both only ever return an answer.
The prompt-level 'resist the case-pack-multiple trap' rule held on this run
0 of 4 planted unrecorded_transfer/case-pack-multiple trap events were drafted as uom_error (uom_transfer_confusion=0).
The live citation-validity check flags a fabricated citation the moment it comes back
src/app.py's /api/draft runs the identical check evals/scoring.py grades the real run with, on every live draft -- confirmed by reading the handler.
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK ON THE CAUSE ITSELF. The rule lives only in src/prompt.py's SYSTEM text -- nothing in src/brief.py or src/app.py recomputes it. See fails.
It does not guarantee an answer at all -- a runaway reasoning pass can consume the entire ceiling and return nothing, which is not the same failure as a wrong cause.
It does not make the drafted cause itself correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against a crafted transaction-history note -- no red-team run exists for this kit (see the security page).
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 19 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
12 measured by the latest run7 need the model half
Cause exact match, plus citation validity, narrative faithfulness and the trap guardrail
alarm
no_answer_events as a raw count, never folded into fabricated_cause_rate_pct without reading which events they are -- see method_note; uom_transfer_confusion specifically, since it is the metric this kit's own eval_intent names -- 0 of 4 this run; answered vs scored on citation validity -- 24 of 36 non-unresolved cause+citation pairs were even in scope; the other 12 correctly said unresolved and cite nothing by design — alarm on Any nonzero uom_transfer_confusion, on any run. Zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
36
different corpus — nothing is comparable
corpus.bytes
35,159
variance events edited — the count held, the bytes did not
split.count
36
the characters per event's assembled header, variance line and transaction log block count moved — a different set was scored
split.size_p50
522.5
the median size of one characters per event's assembled header, variance line and transaction log block moved
split.size_p95
621.0
the 95th-percentile size of one characters per event's assembled header, variance line and transaction log block moved
dataset.rows
36
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
cause accuracy
not yet known
36 events
77.8% this run (28 of 36) -- no repeat, r001-inv-cycle ran once.
citation validity
not yet known
24 non-unresolved cause+citation pairs scored
91.7% this run (22 of 24 non-unresolved pairs scored) -- both scored failures are the 2 no-answer events, not a genuinely fabricated citation; see Eval.method_note. No repeat.
fabricated cause rate (citation-validity axis)
not yet known
24 non-unresolved cause+citation pairs scored
8.3% this run (2 of 24 non-unresolved pairs scored) -- both hits are the 2 no-answer events, not a genuinely fabricated citation on a drafted cause; see Eval.method_note. No repeat.
uom/transfer trap confusion -- this kit's headline guardrail metric
0 -- zero this run
4 planted case-pack-multiple trap events
measured, one run
no-answer rate
2 -- the residual after four MAX_TOKENS increases
36 events
measured, one run per ceiling (2500/4096/8192/16384); never reached zero
35.3% this run (12 of 34 narratives scored) -- a literal substring check, not a semantic one; see Eval.method_note and Business.not_good_enough for why this understates grounded narrative quality. No repeat.
reasoning-token share of output
not yet known
36 calls
98.5% this run (147,503 of 149,762 output tokens), one run, one setting -- reasoning left at provider default, never compared against the live app's thinking-off setting. Narrative only, not a registered metric key.
latency
not yet known
36 calls
p50 13,703ms, p95 145,302ms on r001-inv-cycle -- the p95 IS one of the two no-answer events. One recorded run, not a distribution.
input volume
0 -- fixed by the corpus and the prompt, not the model
36 calls
33,561 input tokens on r001-inv-cycle. Any movement means the prompt or the corpus changed.
output volume
not yet known
36 calls
149,762 output tokens on r001-inv-cycle -- model-specific, and includes whatever reasoning the provider chose to spend, including 32,768 tokens spent on the two no-answer events.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-inv-cycle 2026-08-20
cause accuracy, %
77.8
cause accuracy traceable, %
73.3
cause accuracy unresolved, %
100.0
citation validity, %
91.7
fabricated cause rate, %
8.3
input tokens, whole run
33561
model latency p50 ms
13703.00
model latency p95 ms
145302.00
narrative faithfulness, %
35.3
no answer events
2
output tokens, whole run
149762
uom transfer confusion rate, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 12 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether reasoning is left on or explicitly disabled
not measured -- this run left it on (provider default); the live app always disables it; the two have never been compared on this corpus
reasoning
src/app.py hardcodes thinking=THINKING_OFF; r001-inv-cycle's own top-level thinking field is null (provider default), while every one of its 36 records carries a nonzero reasoning_tokens (300-16,384), and the two no-answer events spent 100% of MAX_TOKENS on it. See Business.not_good_enough.
the planted case-pack-multiple trap (4 of 36 events, an unrecorded_transfer variance that also divides evenly by a common case-pack size)
uom_transfer_confusion on the 4 trap events: 100% (the case-pack-multiple floor, 4 of 4 miscalled uom_error) -> 0% (the fast tier) on the identical events
measured
results/eval-b000-pack-multiple.json vs results/eval-r001-inv-cycle.json, same 4 trap events, one variable (which decider reads the log).
the reasoning-reply-ceiling split
no-answer rate: 17 of 36 (MAX_TOKENS=2500) -> 14 of 36 (4096) -> 5 of 36 (8192) -> 2 of 36 (16384, registered) -- falling sharply but never reaching zero across four real runs
measured
results/eval-r001-inv-cycle.json and the three superseded ceiling attempts (read and independently re-scored before being deleted as scratch evidence; see provenance.verified_by).
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
cause accuracy
nothing yet
citation validity
nothing yet
fabricated cause rate (citation-validity axis)
nothing yet
uom/transfer trap confusion -- this kit's headline guardrail metric
any nonzero value, on any run
no-answer rate
any no-answer event, on any run -- read finish_reason and reasoning_tokens before trusting an accuracy figure computed over that run (see Guardrails.add_first)
any change without a corresponding change to prompt or corpus
output volume
any call reaching MAX_TOKENS (16,384) -- that is a truncated answer, not a long one, and this run had two
NextThe three you would add first
A hard ceiling on the reasoning pass specifically, separate from the total output ceiling, so a runaway reasoning trace fails fast rather than silently consuming the entire MAX_TOKENS budget.Today MAX_TOKENS bounds the total of reasoning plus visible reply and the two are not capped separately, which is mechanically why an event can spend 100% of the ceiling on reasoning and 0% on the answer -- see Business.not_good_enough.
A live, code-level check that recomputes the drafted cause from the log directly (src/segment.py::classify(), the same function that writes gold), and flags any disagreement with the model's own call before the UI renders it.This kit's own scorer (evals/scoring.py) already does this offline, against gold; nothing does it live, against the model's own reply.
A red-team run against the transaction history log's own note text.In a real deployment the log's receiving/transfer/scan/adjustment notes arrive from external systems this kit's architecture treats as trusted with no verification step -- unmeasured for this kit (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting in src/prompt.py or src/app.py -- this kit's own history is that both materially move the no-answer rate.
What this cannot tell you
Whether reasoning explicitly disabled (matching the live app's own setting) reproduces this run's accuracy, latency or no-answer rate -- see Business.not_good_enough.
Whether the no-answer rate is stable across repeated runs at MAX_TOKENS=16384, or would keep falling with a fifth, sixth ceiling increase -- one recorded run at each of four ceilings is not a distribution.
Whether a live, code-level cause check (see add_first) would ever have caught a real disagreement -- this run's own citation check never had a fabrication on a genuinely drafted cause to catch.
Whether the prompt-only trap-resistance rule holds against a transaction-history note crafted to mimic this kit's own flag phrasing -- no red-team run exists for this kit.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is an empty file with a comment explaining that the emptiness is load-bearing. The whole drafting decision is three files: src/prompt.py, src/segment.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
36 variance events, generated from a fixed seed, never fetched. Deriving gold from the same log and variance_qty the model reads (never a true_cause string typed directly) via src/segment.py::classify() -- with tools/verify_gold.py re-deriving every row independently and asserting zero drift -- is what keeps the internal-consistency check honest; see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the five-cause vocabulary, the citation rule and the narrative instruction are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. Passing thinking through is a kwarg on one call site, not a client-library upgrade -- and this kit's own MAX_TOKENS history (Business.not_good_enough) shows why that one kwarg matters more here than on most sibling kits.
evaluation
evals/scoring.py
eval harnesses
exact match on a five-value cause vocabulary, plus citation validity, narrative faithfulness and one named trap-confusion guardrail, is a loop and a handful of counters, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per event -- load, pack, prompt, call, parse -- with no branching and no state carried between events. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A sixth named cause this kit's fixed five-member CAUSE_VOCAB does not carry needs a hand-written src/rubric.py entry plus new flagged evidence lines in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here, which is exactly why this run's own reasoning-token burn had to be read off the raw JSON by hand rather than a dashboard.
No built-in reasoning-budget control beyond the single MAX_TOKENS ceiling -- a framework with a separate reasoning-token cap or a stop-and-retry policy on a runaway trace might have caught this kit's own no-answer events differently; this kit's stdlib-only design did not build one (see Guardrails.add_first).
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-inv-cycle on the fast tier, 2026-08-20. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
13,703 ms
not yet known
nothing yet
Model, p95
145,302 ms
not yet known
nothing yet
Input tokens
33,561
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
149,762
not yet known
any call reaching MAX_TOKENS (16,384) -- that is a truncated answer, not a long one, and this run had two
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — one document is one unit, whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-20, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
variance events
data/events.jsonl -- 36 events, generated once from a fixed seed by tools/build_corpus.py
read by id in src/brief.py and src/app.py; never modified after generation
gold causes
data/gold.jsonl -- 36 rows, computed by tools/build_corpus.py and src/segment.py::classify at generation time, from the same log and variance_qty the model sees
never -- evals/scoring.py is pure code, no model, no key
the key
.env -- never committed
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/draft handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 36 variance events across item/location/period combinations, each with its own 4-8 line transaction history log -- byte-identically from a fixed seed (SEED = 20260819) every time it is run. There is no incremental refresh; src/segment.py::classify()'s flag-precedence rule is read fresh from each event's own log on every request, never baked into the corpus at generation time.
0.046s wall time to regenerate all 36 events and their gold causes, measured 2026-08-20 -- see Data.index. (lenses.Data.index, tools/build_corpus.py)
a real deployment's transaction history changes continuously and grows per item/location over time, not on a fixed seed -- how a genuinely rolling history (the same item recurring cycle over cycle, sometimes still unresolved from last count) would change what this kit can resolve was not measured; see Architecture.breaks_at_scale.
point tools/build_corpus.py at your own transaction types and item master, and every published accuracy/trap figure is void -- they are this corpus's own flag vocabulary and case-pack sizes (see Data.bring_your_own_boundary), not a property of the model.
model
one call per EVENT carrying that item/location's own transaction history log and the variance, behind src/adapters/__init__.py, reasoning left at the provider's documented default (on) -- a DIFFERENT configuration from the one the shipped app's own /api/draft hardcodes (off). See Business.not_good_enough.
34 of 36 events returned a parsed reply (94.4%); 2 of 36 hit finish_reason='length' with reasoning_tokens==output_tokens==16384 and produced no answer at all. (lenses.Eval.scores, r001-inv-cycle)
MAX_TOKENS=16384 was reached and exhausted by exactly the two no-answer events -- not headroom on those two, though the 34 completed calls used a median of 1,175 output tokens, well under it. Whether a sixth ceiling increase would clear the residual is untested; four increases already did not reach zero. See Business.not_good_enough.
the verdict is per-model and one configuration ran once -- reasoning-on is what is scored here, reasoning-off (the shipped app's own setting) has never been run against this corpus; see Eval.could_not_verify.
labels
data/gold.jsonl, 36 rows -- each event's cause, citation and confirmed_note are decided by src/segment.py::classify() from the event's own log and variance_qty, never hand-typed, and tools/verify_gold.py re-derives every row independently and asserts zero drift.
6 gold-unresolved, 9 unrecorded_transfer (4 of them the planted case-pack-multiple trap), 7 mis_receipt, 7 uom_error, 7 unscanned_movement. (lenses.Eval.dataset, inv-cycle-2026-08-19-36events)
your own transaction history: hand-decide which variances are traceable and why, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-adjudicated gold has an error rate this kit has never measured.
cause accuracy and citation validity over this set reflect THIS corpus's own flag vocabulary (receiving_correction / counterpart_activity / uom_note) -- a real deployment's own transaction-log conventions are untested; see Data.bring_your_own_boundary.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
finish_reason='length' on a record, with reasoning_tokens == output_tokens == max_tokens
the entire reply ceiling was spent on reasoning; no cause, citation or narrative was produced at all -- read as 'no answer', never as a wrong judgement.
check finish_reason and reasoning_tokens before trusting any accuracy figure computed over this run -- 2 of 36 records this run, a different 2 than the 5/14/17 events that showed the same symptom at each smaller ceiling tried (8192/4096/2500). (results/eval-r001-inv-cycle.json, records VC-0010 and VC-0019)
a drafted cause of unresolved on a gold-traceable event
either a correct, honest miss -- the log genuinely lacks a clean signal -- or the model failed to reconcile the log's own arithmetic against the stated variance, most visibly on unscanned_movement (6 of the model's 12 unresolved calls this run are misses on that one gold class, 6 of 7 gold unscanned_movement events total).
check whether gold's own cause is unscanned_movement before reading an unresolved call as evidence the log has nothing to say. (results/eval-r001-inv-cycle.json per_event rows where model_cause='unresolved' and gold_cause != 'unresolved')
citation_valid=False on a per_event row
either a genuinely fabricated citation (a real line cited for a cause it does not support) or a no-answer event whose null cause is scored as an automatic citation failure by convention -- see Eval.method_note. This run's own 2 hits are entirely the latter.
check whether model_cause is null before reading a citation_valid=False row as evidence of fabrication. (results/eval-r001-inv-cycle.json, fabricated_examples)
No machine symptom — this failure leaves no trace in any output.
narrative_faithfulness() is a literal substring check on the cause's own label text, with no semantic fallback -- a narrative that correctly grounds the drafted cause in plain prose ('a unit-of-measure mismatch' rather than the token 'uom') scores identically to one that never engaged with the cause at all. Both render as narrative_faithful=False, and nothing in the record distinguishes a substantively-faithful narrative from a substantively empty one beyond reading the text by hand -- which is what this build session did for the 14 non-unresolved 'unfaithful' hits (see Business.not_good_enough), not something the pipeline itself checks.
Concurrency and GPU sizing -- one serial call per event, nothing measured past 36. Provider-side retention, training use and log residency -- provider-dependent, and this kit's payload is a transaction history log naming item and location codes, so that unknown IS the posture question. Whether the provider's prompt cache was hit on any call -- Cost prices every call at the cache-miss rate for exactly this reason. Whether the no-answer rate is a stable per-event probability or an artifact of this specific 36-event sample -- four ceilings were tried and none reached zero, and a fifth was deliberately not fired; see Business.not_good_enough. Whether the same figures hold with reasoning explicitly disabled, matching the live app's own setting.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Cause exact match, plus citation validity, narrative faithfulness and the trap guardrail
Find why a store's stock count doesn't match the system
PresenterOpens the private repo. Visible to admins only.
In one lineCause exact match, plus citation validity, narrative faithfulness and the trap guardrail
Does the drafted cause match gold's five-way vocabulary exactly, including saying unresolved when gold says unresolved? For a non-unresolved cause, is the cited log line real AND does it actually support that cause? Does the narrative's own prose name the drafted cause? Of the 4 planted case-pack-multiple trap events, how many were drafted as uom_error instead of the gold unrecorded_transfer?
$0.00per 1,000 variance events
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Event
VC-0004
Item
SKU-75000
Item label
Batteries
Location
LOC-007
Variance (system minus counted)
-144
Planted scenario
unrecorded_transfer_trap
Log line 0 (the citation both gold and the model use)
[adjustment] Regional log shows LOC-022 posted a partial adjustment of 144 units around this same count window -- not logged here as a transfer.
Gold cause
unrecorded_transfer
Gold citations
[0]
The case-pack-multiple floor's call
uom_error (WRONG -- 144 divides evenly by both case-pack sizes, 12 and 24, so the floor calls it uom_error without checking any log line)
The model's call
unrecorded_transfer
The model's citations
[0]
The model's drafted narrative
The cited regional log shows LOC-022 posted a 144-unit partial adjustment in the same count window, but no corresponding transfer was logged at LOC-007. This matches the 144-unit variance and indicates an unrecorded transfer between locations.
Scored as
correct -- the trap the free case-pack-multiple floor fails and the fast tier does not
Grader
Verdict
Why
Cause exact match, plus citation validity, narrative faithfulness and the trap guardrail
correct
VC-0004: item SKU-75000 (Batteries) at LOC-007, variance -144, divisible by both this corpus's case-pack sizes (12 and 24). Log line 0 reads "Regional log shows LOC-022 posted a partial adjustment of 144 units around this same count window -- not logged here as a transfer." Gold: unrecorded_transfer, citing line 0, is_trap=true. The case-pack-multiple floor sees 144 divides evenly by 12 and calls it uom_error without checking any log line -- wrong, by construction. The fast tier read the log, correctly drafted unrecorded_transfer citing line 0, and its narrative named exactly the regional adjustment the citation shows and nothing else.
The formulaWhat it computes
cause_accuracy = exact matches / 36. citation_validity = valid citations / (non-unresolved causes scored, where a falsy cause also counts as an automatic failure). narrative_faithfulness = narratives containing the cause label / narratives produced. uom_transfer_confusion = trap events drafted as uom_error / 4 planted trap events.
The analysisWhat it actually did
Model
Result
the fast tier
77.8% field accuracy
In operationWhat to monitor
Reference standard: tools/build_corpus.py plants a flagged log line per scenario (receiving_correction, counterpart_activity, uom_note) or leaves none for unresolved; src/segment.py::classify() is the only thing that turns that into gold's cause and citations, and tools/verify_gold.py re-derives every row independently and asserts zero drift.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own flag placement, which comes from tools/build_corpus.py's fixed template set.
Watch these
no_answer_events as a raw count, never folded into fabricated_cause_rate_pct without reading which events they are -- see method_note
uom_transfer_confusion specifically, since it is the metric this kit's own eval_intent names -- 0 of 4 this run
answered vs scored on citation validity -- 24 of 36 non-unresolved cause+citation pairs were even in scope; the other 12 correctly said unresolved and cite nothing by design
Alarm on
Any nonzero uom_transfer_confusion, on any run. Zero this run.
How tight can the band be? There is no tolerance band -- cause is an exact match over a five-value vocabulary, and citation validity is a binary real-and-supporting check. Nothing here is a continuous quantity to round.
Cadence: Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting.
The decisionWhen to reach for it
Use it
Gold is DERIVED from the same generated log and variance_qty the model reads -- true of every kit corpus, never true of a real inventory-control team's own variance history.
Do not use it
The true cause is not known in advance -- the normal state of a real cycle-count program, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning