Explain a retailer's month-end inventory change, line by line
Each month, an item's inventory value moves, and the feed behind it is a list of unlabelled lines. This app sorts each line into receipts, markdowns, shrink, allowances, transfers or unknown, totals them, and sizes what is left against your threshold.
PresenterOpens the private repo. Visible to admins only.
For cost accountingCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
Cost accounting at a retailer, explaining each item's inventory value by warehouse and month before the books close.
✕Today's manual process
1Read every line in the month's feed and decide: receipt, markdown, shrink, allowance or transfer.
2Total each group in a spreadsheet, and work out what the lines leave unexplained.
3Compare the gap to the account's threshold, for every item and location, every close.
4One misread line puts a transfer in shrink, and two totals are wrong at once.
Every line sorted and totalled manually
✓With the app
1Every line is sorted into one of the five groups, or marked unknown when the memo does not say.
2Each group is totalled, and the gap the lines leave unexplained is sized.
3The gap is checked against your threshold and marked within it or over it.
4Cost accounting reviews the draft and decides. The app books no adjustment and closes nothing.
People review a sorted, totalled draft
See it work
One real case: what the app reads, step by step
Insulated lunch bags at the north warehouse in June: 17 lines, value down from $56,248.78 to $53,301.33, threshold $750.
Explain a retailer's month-end inventory change, line by lineReference appBuilt to be shaped to your process
5
1One item, one month insulated lunch bags at the north warehouse, June.
2The change to explain value fell from $56,248.78 to $53,301.33; the threshold is $750.
3A look-alike line it opens like a transfer, but the memo ends: written off as shrink.
4The app's call not a transfer: the same line, sorted, filed as shrink.
5Not fully explained the gap is above the $750 threshold; the app never calls that fully explained.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Explain a retailer's month-end inventory change, line by line
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Reconciling one item/location/period's inventory roll-forward means taking every raw, unclassified transaction line for the period and sorting it by hand into receipts, markdowns, shrink write-offs, vendor allowances or inter-location transfers, summing each bucket to a dollar total, and sizing whatever residual is left against the account's stated materiality threshold -- for every cycle, every close, with a transfer-out line and a shrink write-off sometimes reading almost identically in the raw feed. Cost accounting reading every raw, unclassified transaction line for one item/location/period roll-forward by hand -- sorting each into receipts, markdowns, shrink, allowances or transfers, summing each bucket, and sizing whatever residual is left against the account's materiality threshold -- for every cycle, every close.
Audience
Cost accounting staff who review inventory roll-forwards before close, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual roll-forward cycles
The corpus is 34 roll-forward cycles, 0.10 MB (jsonl 1). A real inventory valuation roll-forward names real SKUs, real store/DC locations and real dollar movements a retailer would not let leave the building -- this repo has never held one and is not able to. There is no public corpus of paired (raw unclassified transaction feed, cost-accounting-reconciled bucket decomposition) records, for the same reason no public corpus of month-end close packages or supplier agreements exists: the interesting cases are exactly the ones nobody can publish. Nothing in the corpus refers to a real company -- item names, location codes, transaction references and memo narratives are all invented.
The corpus
The 34 roll-forward cyclesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your roll-forward cycles. That is the whole change — there is no database to migrate.
One roll-forward cycle, as the model receives itcycles.jsonl · 1 of 34
{"cycle_id": "IVR-00001", "item_id": "SKU-14588", "item_label": "Insulated Lunch Bag", "location_id": "DC-NORTH", "period": "2026-06", "opening_balance": 56248.78, "closing_balance": 53301.33, "materiality_threshold_usd": 750.0, "lines": [{"line_id": "IVR-00001-L01", "date": "2026-06-19", "qty": 305, "amount_usd": 3003.27, "memo": "Inbound shipment received against purchase order PO-59265 (ref TXN-10001)."}, {"line_id": "IVR-00001-L10", "date": "2026-06-24", "qty": 46, "amount_usd": -473.69, "memo": "Return-to-vendor allowance posted against inventory (ref TXN-10010)."}, {"line_id": "IVR-00001-L13", "date": "2026-06-28", "qty": 31, "amount_usd": 76.05, "memo": "Inbound transfer received from STR-221 (ref TXN-10013)."}, {"line_id": "IVR-00001-L07", "date": "2026-06-08", "qty": 69, "amount_usd": -954.22, "memo": "Units removed from location, ref TXN-10007. Reason: cycle count discrepancy, written off as shrink."}, {"line_id": "IVR-00001-L17", "date": "2026-06-02", "qty": null, "amount_usd": 1316.88, "memo": "Adjustment posted per ticket reference, ref TXN-10017. No further detail on file."}, {"line_id": "IVR-00001-L08", "date": "2026-06-23", "qty": 47, "amount_usd": -840.21, "memo": "Physical inventory count shortage at this location, written off as shrink (ref TXN-10008)."}, {"line_id": "IVR-00001-L11", "date": "2026-06-11", "qty": 130, "amount_usd": -446.24, "memo": "Vendor damage allowance credited against inventory value (ref TXN-10011)."}, {"line_id": "IVR-00001-L06", "date": "2026-06-18", "qty": 73, "amount_usd": -794.83, "memo": "Units removed from location, ref TXN-10006. Reason: cycle count discrepancy, written off as shrink."}, {"line_id": "IVR-00001-L14", "date": "2026-06-04", "qty": 246, "amount_usd": 114.9, "memo": "Inbound transfer received from STR-517 (ref
Abridged — the file continues.
The outcomeWhat a good result looks like
A bucket per raw line (receipts, markdowns, shrink, allowances, transfers, or an honest unknown when the memo genuinely does not support any of the five), the five bucket totals, and a residual sized and characterized against the account's stated materiality threshold -- for cost accounting to review before close. Decision-free by design: this kit never books an inventory adjustment or closes the valuation period itself.
And when it cannot
A decomposition a preparer trusts that is not real -- a transfer-out line silently read as a shrink write-off (or the reverse), which misstates both bucket totals at once, or a real, material residual characterized as within_threshold, the expensive-direction error false_fully_explained exists to catch. Measured at 0 of 34 cycles and 0 of 588 lines this run -- see not_good_enough for why zero on one clean-language run is not the same claim as zero on a harder or adversarial corpus.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Bucketing a period's raw transaction lines and sizing the residual for cost accounting to pre-check before close — the fast tier, over the free keyword floor 100% line accuracy and 0% transfer<->shrink confusion this run, against the floor's 94.7% and 26.3% confusion on the exact lines written to be confusable.
Deciding whether a real, above-threshold inventory movement gets escalated — the fast tier's residual_status verdict, read alongside its bucket totals 0 of 13 gold above-threshold cycles this run were called within_threshold (false_fully_explained) -- the expensive-direction error this kit's guardrail exists to catch.
What this kit is not — a draft for cost accounting to review before close -- never an approval src/invval.py and src/app.py have no function anywhere that books an inventory adjustment or marks the valuation period closed; every decomposition is for a human to act on.
At a glanceHow the whole thing runs
100%line accuracy pct
14,306 msp50, end to end
$7.58per 1,000 roll-forward cycles · Google Gemini 3 Flash
Run once, for real, on 2026-08-19. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Explain a retailer's month-end inventory change, line by line14 steps · 4 questions · run once, for real · 2026-08-19
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own items, locations and transaction-memo phrasing, or write your own data/cycles.jsonl in the same shape -- src/invval.py and src/prompt.py read a cycle by its own fields (opening_balance, closing_balance, materiality_threshold_usd, lines[]) and do not care where the numbers came from. The measured 100% line accuracy, 0% transfer<->shrink confusion and 100% residual-status accuracy are THIS corpus's memo phrasing variety (a handful of hand-written templates per bucket, with 45% of shrink and transfer-out lines deliberately written in the confusable shared opening) and THIS corpus's one-defect-family-per-cycle design (a cycle's residual is exactly the signed sum of its unknown lines, never a genuinely disputed judgment call).Corpus lens →
When is this the wrong choice?
Avoid: The keyword floor for anything beyond receipts/markdowns/allowances/unknown, where its fixed rules happen to match this corpus's phrasing -- it is not a competitor, it is the honest floor a model has to clear. That is the case against the best-fitting scenario (“Bucketing a period's raw transaction lines and sizing the residual for cost accounting to pre-check before close”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
This kit assumes every feed (receipts, markdowns, shrink, transfers) posts at the SAME item/location/period grain it reads at -- a real deployment where one feed only nets at the department level cannot be decomposed cleanly by this kit; its lines would arrive already netted across items this corpus never mixes, and no amount of careful reading recovers the per-item split from a departmental total. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether this run's 100% line accuracy, 0% transfer<->shrink confusion and 0 false_fully_explained are stable properties of this model at these settings, or a one-off -- one recorded run is not a history. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-19 — r001-fin-invval. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- every roll-forward cycle and its raw lines -- since data/cycles.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Clicking Decompose with no API_KEY returns a calm 200 explaining nothing was called (src/app.py's /api/check). It cannot reproduce a live decomposition, a score, or a dollar figure without a key -- those are what results/eval-r001-fin-invval.json already committed.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
14,306 msp50, end to end
21,336 msp95
1 minclone to first result
What the clock covers. END-TO-END per roll-forward cycle: one HTTP request carrying the cycle's opening balance, closing balance, materiality threshold and every raw transaction line, parsed to a bucket per line plus the bucket totals and residual. No retrieval step -- loading a cycle from disk is pure code and costs no network time. This run left provider-side reasoning ('thinking') at its documented default (on) rather than disabling it, so p50/p95 include a hidden reasoning pass on every call, not only the JSON-formatting cost -- see not_good_enough.
Current processWhat it replaces
Cost accounting reading every raw, unclassified transaction line for one item/location/period roll-forward by hand -- sorting each into receipts, markdowns, shrink, allowances or transfers, summing each bucket, and sizing whatever residual is left against the account's materiality threshold -- for every cycle, every close.
Where it is not good enough
Two real findings. First, a live discrepancy between what this report measures and what a forker's own click actually does: the registered run (r001-fin-invval) left provider-side reasoning ('thinking') at its documented default -- ON -- while src/app.py's /api/check hardcodes thinking=THINKING_OFF on every live call. Nobody reconciled the two before this run was registered as the measured fact. Reasoning consumed 79.3% of this run's average output-token budget (60,159 of 75,867 total output tokens; 1,769.4 tokens/call average, up to 3,154 on one cycle) and is the largest single driver of both this kit's latency (p50 14.3s, p95 21.3s -- roughly 6-10x fin-close's 2.2s/2.5s on the same model with reasoning explicitly disabled) and, since a provider bills reasoning tokens as completion tokens, its dollar cost: reasoning tokens are already inside the 75,867 output tokens cost_per_query_usd is computed from. The published cost and latency numbers on this report are therefore NOT what a reader gets from the shipped UI -- that configuration has not been separately measured. Second, a structural limitation stated in the corpus's own documentation, not found by this run: data/SOURCES.md states plainly that this kit assumes every feed (receipts, markdowns, shrink, transfers) posts at the SAME item/location/period grain it reads at -- a real feed that only nets at the department level cannot be decomposed cleanly by this kit at all, no matter how carefully the memo is read. Separately, this run's own 100% line accuracy, 0% transfer<->shrink confusion and 0 false_fully_explained are measured on one 34-cycle, 588-line run with clean, non-adversarial memo language and exactly one planted defect family per cycle -- not yet a distribution (no repeat), and not a claim about a corpus written to actively mislead the model (no red-team run exists for this kit; see Eval.could_not_verify).
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free keyword baseline (evals/baseline.py, a fixed-priority rule that checks 'removed from location' before 'transfer') scores 94.7 pct line accuracy overall but misclassifies 31 of 118 gold transfer lines as shrink — exactly the deliberately confusable phrasing this kit's model classified correctly on all 75 planted instances. No red-team run exists for this kit — unlike fin-close's own figure, this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
whether reasoning is enabled
src/invval.py and src/app.py
check() takes a thinking kwarg; src/app.py's live /api/check hardcodes it OFF while the registered eval run (r001-fin-invval) left it at provider default (on) -- a real, live discrepancy between what this report measures and what a forker's own click experiences today. See Business.not_good_enough.
the bucket taxonomy
src/prompt.py
A bucket this kit's fixed five-member list does not carry needs a new entry in BUCKETS/BUCKET_MEANINGS here plus new template memo phrasing in tools/build_corpus.py -- a prompt tweak alone will not teach a label the corpus never plants.
the corpus
tools/build_corpus.py
Point it at your own items, locations and transaction-memo phrasing. src/invval.py and src/prompt.py read a cycle by its own fields (opening_balance, closing_balance, materiality_threshold_usd, lines[]) and do not care where the numbers came from.
the materiality threshold
data/cycles.jsonl (materiality_threshold_usd per cycle)
Set per cycle by the corpus, not a single global constant to edit -- a new corpus states its own threshold per account/cycle in the same field.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 34 roll-forward cycles and 588 raw transaction lines from a fixed seed (SEED=20260819). Gold (a bucket per line, bucket totals, residual, materiality status) is RECOMPUTED from the actual generated lines, never from the random targets that seeded them -- an internal-consistency check (_verify) asserts every cycle reconciles to the cent before the build is considered good.
the prompt
src/prompt.py
The five-bucket-plus-unknown vocabulary, the sign convention (receipts/markdowns/shrink/allowances positive magnitudes, transfers a signed net figure) and the answer schema, declared once and read from here by build() and parse(). States the transfer-vs-shrink confusion to the model explicitly, as the named failure mode this task exists to catch, not left for the model to discover.
the AI layer
src/invval.py
Loads a cycle by id, calls the model once with every raw line in the cycle batched into that single call, parses the reply into a bucket per line plus bucket totals and the residual. The whole AI layer, deliberately short -- the same split fin-close's close.py and gap-brief's brief.py both make. Never books an adjustment or closes the valuation period -- there is no function here or in src/app.py that does either.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
spend control
src/budget.py
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
the app
src/app.py
Static files and two JSON endpoints (/api/cycle, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-fin-invval) actually used; see Business.not_good_enough.
the scorer
evals/scoring.py
Line-level and cycle-level scoring, pure code: exact match per line against gold, transfer<->shrink confusion reported as its own named figure rather than folded into a generic 'wrong' bucket, and false_fully_explained (a real, above-threshold residual called within_threshold) counted separately as the expensive-direction error.
free baseline
evals/baseline.py
A fixed-priority keyword rule, written the way a person free-texting a quick classifier would write one, that deliberately fires on 'removed from location' BEFORE the 'transfer' keyword check -- so it misclassifies every ambiguous transfer-out line as shrink by construction. The honest, narrow floor a model has to clear.
Where it breaks at scale
Not on cycle count -- each cycle is independent and one call per cycle is linear. It breaks two other ways. First, GRAIN: this kit assumes every feed (receipts, markdowns, shrink, transfers) posts at the SAME item/location/period grain it reads at (see Data.breaks_on) -- a real feed that only nets at the department level cannot be decomposed cleanly by this kit at all, no matter how carefully the memo is read; this is stated plainly in data/SOURCES.md, not discovered by this run. Second, LINE COUNT PER CYCLE: MAX_TOKENS is fixed at 4096 (src/invval.py), sized from this run's own measurement -- the largest cycle here is 23 raw lines, and the worst observed cost was roughly 126 tokens/line including reasoning, with headroom to spare. A cycle with meaningfully more than ~23 raw lines was never measured and would need a larger MAX_TOKENS or a split-and-merge strategy, since one call still carries every line in a cycle at once by design (see LLM lens). This run also left reasoning at the provider's default (on) rather than disabling it explicitly -- reasoning consumed 79.3% of this run's average output-token budget (1,769.4 of 2,231.4 tokens/call) -- so a cycle near the MAX_TOKENS ceiling is far more exposed to a reasoning-driven truncation than a non-reasoning run of the same size: the very first attempt at this run, at MAX_TOKENS=1800, hit finish_reason='length' on 24 of 34 cycles for exactly this reason (src/invval.py's own comment).
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Cycle IVR-00001 -- SKU-14588 (Insulated Lunch Bag), DC-NORTH, period 2026-06: 17 raw transaction lines shown against a $56,248.78 opening balance, $53,301.33 closing balance and a $750.00 materiality threshold, before any call is made.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page after pressing Decompose with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. No decomposition is shown -- this is the honest failure state, not a staged one.failureOpen full size →
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
34roll-forward cycles
0.10 MiBjsonl 1
p50 17chars per raw transaction lines per cycle
$0.00setup · 0.0s
How it is cutWhat one raw transaction lines per cycle is
One cycle, every raw line included whole -- no train/test split, every cycle is decomposed once, in one call.
SetupWhat the setup figure measured
No index is built. src/app.py finds a cycle by scanning C.cycles() for a matching cycle_id -- 34 rows, a linear scan, not a search structure -- and src/invval.py sends that cycle's raw lines whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own roll-forward cycles
Point tools/build_corpus.py at your own items, locations and transaction-memo phrasing, or write your own data/cycles.jsonl in the same shape -- src/invval.py and src/prompt.py read a cycle by its own fields (opening_balance, closing_balance, materiality_threshold_usd, lines[]) and do not care where the numbers came from. The five-bucket taxonomy in src/prompt.py assumes receipts/markdowns/shrink/allowances/transfers specifically; a different bucket set is a prompt.py change, not a data change.
⚠︎ And what stops being true when you do: The measured 100% line accuracy, 0% transfer<->shrink confusion and 100% residual-status accuracy are THIS corpus's memo phrasing variety (a handful of hand-written templates per bucket, with 45% of shrink and transfer-out lines deliberately written in the confusable shared opening) and THIS corpus's one-defect-family-per-cycle design (a cycle's residual is exactly the signed sum of its unknown lines, never a genuinely disputed judgment call). A real transaction feed's own memo language, and a feed that does not post at a matching item/location grain (see breaks_on), are both unmeasured by this run.
What breaks it
This kit assumes every feed (receipts, markdowns, shrink, transfers) posts at the SAME item/location/period grain it reads at -- a real deployment where one feed only nets at the department level cannot be decomposed cleanly by this kit; its lines would arrive already netted across items this corpus never mixes, and no amount of careful reading recovers the per-item split from a departmental total.
One planted ambiguity family, transfer-out vs. shrink. A real roll-forward carries confusions this corpus does not attempt -- a vendor allowance that is partly a shrink write-off, a transfer corrected mid-period, a receipt split across two GL postings.
The residual check is a threshold comparison, not a judgment call. A real cost accountant sometimes accepts an above-threshold residual with a documented explanation found outside the raw feed, or escalates a below-threshold one anyway; this corpus's gold residual status is derived mechanically from the materiality threshold alone.
One cycle, one static snapshot -- not a rolling history. A real deployment would see the same item/location recur period over period; this corpus judges every cycle independently.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,269
1,012
cycle
176
55
lines
2,209
684
Total
1,751
This is the cost lesson as arithmetic: of the 1,751 tokens assembled, 1,012 are instructions — 58% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged verbatim by importing src.prompt and src.invval directly and calling prompt.build(cycle) for IVR-00001 (not retyped) -- the literal system and user content src/invval.py::check() sends. Per-part token counts are a proportional estimate over character share (system 3269 / cycle 176 / lines 2209 chars, 5654 total) applied to this call's own total of 1751 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are drafting an inventory valuation roll-forward decomposition for one item/location/period, for cost accounting to review before close. You are given the opening balance, the closing balance (both from the perpetual/GL system -- trusted, not disputed), a materiality threshold in dollars, and every RAW, UNCLASSIFIED transaction line for the period.
For EACH raw line, classify it into exactly one of these five buckets, or `unknown` if the memo genuinely does not support any of them:
receipts inbound stock received against a purchase order or vendor shipment. Increases the balance.
markdowns a permanent price/value reduction posted against inventory (clearance, promotional revaluation). Decreases the balance.
shrink a write-off from a cycle count variance, damage, spoilage or loss -- inventory leaves the books with no receiving location on the other end. Decreases the balance.
allowances a vendor damage/return allowance or similar credit posted against inventory value, not a physical unit movement. Decreases the balance.
transfers a movement of the SAME item to or from another location -- transfers-in increase the balance, transfers-out decrease it. Always names or implies a counterparty location or a transfer reference.
unknown the memo does not name or clearly support any of the five buckets above. Do not force a line into a bucket to avoid saying unknown -- an unsupported guess is graded worse than an honest unknown.
⚠︎ THE FAILURE MODE THIS TASK EXISTS TO CATCH: a transfer-out line and a shrink write-off line can open with nearly identical language -- both are often phrased as "units removed from location". Do not classify from the opening clause alone. A transfer names or implies a counterparty location and/or a transfer reference; a shrink write-off states a count variance, damage or loss reason and no destination. Read the WHOLE memo before deciding between these two.
SIGN CONVENTION -- state your bucket totals this way, exactly:
receipts, markdowns, shrink and allowances are each a POSITIVE dollar magnitude (the size of that movement, regardless of direction).
transfers is a SIGNED net figure: transfers-in positive, transfers-out negative, netted together.
receipts and transfers-in increase the balance; markdowns, shrink, allowances and transfers-out decrease it.
Sum each bucket's classified lines to that bucket's total (lines you called `unknown` are excluded from every bucket total). Then compute:
residual = (closing_balance - opening_balance) - (receipts - markdowns - shrink - allowances + transfers)
State whether the residual is within or above the stated materiality threshold, using the absolute value of the residual. An above-threshold residual is NEVER `within_threshold` -- do not round it down or describe it as fully explained. If every line is confidently classified and the residual is still above threshold, say so: that is a real unexplained movement for cost accounting to chase, not a rounding error to wave through.
You are drafting this decomposition for cost accounting to review. You never book an inventory adjustment and you never close the valuation period yourself -- you only classify, total and size the residual for a human to act on.
CYCLE IVR-00001 -- item SKU-14588 (Insulated Lunch Bag), location DC-NORTH, period 2026-06
Opening balance: $56248.78
Closing balance: $53301.33
Materiality threshold: $750.00
RAW TRANSACTION LINES (17)
--------------------------
- IVR-00001-L01 | 2026-06-19 | qty 305 | $3003.27 | Inbound shipment received against purchase order PO-59265 (ref TXN-10001).
- IVR-00001-L10 | 2026-06-24 | qty 46 | $-473.69 | Return-to-vendor allowance posted against inventory (ref TXN-10010).
- IVR-00001-L13 | 2026-06-28 | qty 31 | $76.05 | Inbound transfer received from STR-221 (ref TXN-10013).
- IVR-00001-L07 | 2026-06-08 | qty 69 | $-954.22 | Units removed from location, ref TXN-10007. Reason: cycle count discrepancy, written off as shrink.
- IVR-00001-L17 | 2026-06-02 | qty n/a | $1316.88 | Adjustment posted per ticket reference, ref TXN-10017. No further detail on file.
- IVR-00001-L08 | 2026-06-23 | qty 47 | $-840.21 | Physical inventory count shortage at this location, written off as shrink (ref TXN-10008).
- IVR-00001-L11 | 2026-06-11 | qty 130 | $-446.24 | Vendor damage allowance credited against inventory value (ref TXN-10011).
- IVR-00001-L06 | 2026-06-18 | qty 73 | $-794.83 | Units removed from location, ref TXN-10006. Reason: cycle count discrepancy, written off as shrink.
- IVR-00001-L14 | 2026-06-04 | qty 246 | $114.90 | Inbound transfer received from STR-517 (ref TXN-10014).
- IVR-00001-L04 | 2026-06-10 | qty 113 | $-1705.82 | Clearance markdown, unit value reduced and inventory revalued (ref TXN-10004).
- IVR-00001-L15 | 2026-06-03 | qty -175 | $-8711.25 | Transfer-out to STR-517 (ref TXN-10015).
- IVR-00001-L02 | 2026-06-17 | qty 220 | $7023.89 | Vendor receipt posted, purchase order PO-43736 (ref TXN-10002).
- IVR-00001-L16 | 2026-06-03 | qty 47 | $1262.29 | Adjustment posted per ticket reference, ref TXN-10016. No further detail on file.
- IVR-00001-L05 | 2026-06-25 | qty 45 | $-590.61 | Inventory removed from floor stock, ref TXN-10005 -- count adjustment, expensed to shrink.
- IVR-00001-L12 | 2026-06-21 | qty 116 | $67.23 | Inter-location transfer, units received from STR-221 (ref TXN-10012).
- IVR-00001-L09 | 2026-06-13 | qty n/a | $-455.78 | Vendor chargeback allowance applied to inventory value (ref TXN-10009).
- IVR-00001-L03 | 2026-06-10 | qty 143 | $-839.31 | Clearance markdown, unit value reduced and inventory revalued (ref TXN-10003).
Return a JSON object with these keys:
"lines": a list with exactly 17 entries, one per raw line above in the same order, each {"line_id": <id>, "bucket": <one of: receipts, markdowns, shrink, allowances, transfers, unknown>}.
"bucket_totals": {"receipts": <number>, "markdowns": <number>, "shrink": <number>, "allowances": <number>, "transfers": <number>} -- per the sign convention stated above.
"residual_usd": <number, the signed residual computed per the formula above>.
"residual_status": <one of: within_threshold, above_threshold>.
Answer with the JSON object only.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Explain a retailer's month-end inventory change, line by line — 34 roll-forward cycles. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code, two levels: LINE level (each raw line's predicted bucket compared exactly to gold, including 'unknown'; transfer<->shrink confusion reported as its own named figure rather than folded into a generic 'wrong') and CYCLE level (residual_status compared exactly to gold's threshold characterization; residual_usd compared to gold within a tolerance band of max($50, 5% of |gold residual|); false_fully_explained -- a gold above_threshold residual called within_threshold -- and false_alarm_material -- the reverse -- counted separately as the two directional errors, never folded into one accuracy number). No judge model. The same function scores both evals/baseline.py's free floor and evals/run.py's real run.
34roll-forward cycles
34source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED588 / 588line accuracy pct — raw lines answeredDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED588 / 588lines answered pct — raw lines askedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED34 / 34residual status accuracy pct — roll-forward cyclesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED34 / 34residual amount correct pct — roll-forward cyclesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 13false fully explained — gold above-threshold cyclesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 21false alarm material — gold within-threshold cyclesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 118transfer as shrink rate pct — gold transfer linesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 106shrink as transfer rate pct — gold shrink linesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/scoring.py) is pure code, exact match against a mechanically-derived gold bucket -- there is no judgement to validate, only arithmetic. What WAS validated: tools/build_corpus.py's gold is RECOMPUTED from the actual generated lines after they are split and rounded to the cent -- never carried over from the random targets that seeded the split -- and an internal-consistency check (_verify) asserts every cycle's bucket totals, residual and closing balance reconcile to the cent before the corpus build is considered good, so a label cannot drift from the record it describes.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One roll-forward cycle
1,000 roll-forward cycles
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.007578
$7.58
12%
Same work, 1× the bill
The same roll-forward cycles, the same tokens — only the rate card changed. And on that card about 12% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether reasoning ('thinking') is left on or explicitly disabled for this model -- the app's own /api/check always disables it; the registered run left it at provider default. The two configurations have not been priced against each other on this corpus.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page. Reasoning was also left at provider default (on) for this run -- see Business.not_good_enough -- so even the projected figure prices tokens the shipped app's own reasoning-off configuration would not spend.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing -- exact match per line and a tolerance-banded residual comparison are both pure arithmetic. A measured $0.00, not an unpriced one.
The gradersOne way to grade, and why it is the only one
The floor's transfers accuracy alone is 73.7% (87 of 118), because its 'removed from location' -> shrink rule is checked BEFORE its 'transfer' keyword rule -- by construction it misclassifies 31 of the 118 gold transfer lines as shrink, and those 31 are exactly the confusable-phrasing transfer-out lines tools/build_corpus.py plants (see Data.corpus and taxonomy). Its residual-status accuracy still reaches 100% because a cycle's residual is the sum of its UNKNOWN lines specifically, not its transfer/shrink split -- so a transfer<->shrink swap can leave the aggregate residual untouched even while misstating both bucket totals. The fast tier's 0% confusion on the identical 118 lines is the real, measured gap a model closes that this floor cannot.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Per-line bucket exact match, plus a tolerance-banded residual and threshold check Does the model's bucket for each raw line match gold (including 'unknown')? Does the cycle's residual dollar figure fall within a tolerance of gold? Does the cycle's above/within-threshold characterization match gold exactly, with the two directional errors (false_fully_explained, false_alarm_material) counted on their own?
$0.00
no
yes
the fast tier 100.0% line accuracy · 1 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, and cleanly on the one axis this corpus was built to test: the fast tier and the free keyword floor tie on receipts, markdowns, allowances and unknown (all 100% for both), then diverge sharply on transfers -- 100% (fast tier) vs 73.7% (floor), because the floor's opening-clause keyword rule misclassifies exactly the 31 gold transfer lines written in the deliberately confusable shared phrasing (see Data.corpus and taxonomy) as shrink, and the fast tier does not misclassify any of them. A grader that could not tell the two apart would not produce a gap that lines up this precisely with the corpus's own documented ambiguous-line counts.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Bucketing a period's raw transaction lines and sizing the residual for cost accounting to pre-check before close
the fast tier, over the free keyword floor
100% line accuracy and 0% transfer<->shrink confusion this run, against the floor's 94.7% and 26.3% confusion on the exact lines written to be confusable.
the keyword floor for anything beyond receipts/markdowns/allowances/unknown, where its fixed rules happen to match this corpus's phrasing -- it is not a competitor, it is the honest floor a model has to clear.
Deciding whether a real, above-threshold inventory movement gets escalated
the fast tier's residual_status verdict, read alongside its bucket totals
0 of 13 gold above-threshold cycles this run were called within_threshold (false_fully_explained) -- the expensive-direction error this kit's guardrail exists to catch.
trusting a within_threshold verdict without checking whether reasoning was left on for the call that produced it -- see Business.not_good_enough; this run's own configuration has not been re-measured with reasoning explicitly disabled the way the live app runs it.
What this kit is not
a draft for cost accounting to review before close -- never an approval
src/invval.py and src/app.py have no function anywhere that books an inventory adjustment or marks the valuation period closed; every decomposition is for a human to act on.
wiring this kit's output straight into a posting or period-close system with no human review step in between.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
receipts
Inbound stock received against a PO or vendor shipment
117
IVR-00001-L01: 'Inbound shipment received against purchase order PO-59265 (ref TXN-10001).' Gold and model both: receipts.
markdowns
A permanent price/value reduction posted against inventory
111
IVR-00001-L04: 'Clearance markdown, unit value reduced and inventory revalued (ref TXN-10004).' Gold and model both: markdowns.
shrink
A count-variance, damage or loss write-off with no receiving location
106
IVR-00001-L07: 'Units removed from location, ref TXN-10007. Reason: cycle count discrepancy, written off as shrink.' Written in the confusable shared opening ('removed from location') on purpose -- the trailing reason clause, not the opening, is what decides…
allowances
A vendor damage/return credit posted against inventory value
105
IVR-00001-L10: 'Return-to-vendor allowance posted against inventory (ref TXN-10010).' Gold and model both: allowances.
transfers
A movement of the same item to or from another location
118
IVR-00003-L19: 'Units removed from location, ref TXN-10054. Destination STR-517, per inter-location transfer directive.' Also written in the confusable shared opening -- the named destination and transfer reference decide it, not the opening clause. Gold and…
unknown
The memo genuinely does not name or support any of the five buckets
31
IVR-00001-L17: 'Adjustment posted per ticket reference, ref TXN-10017. No further detail on file.' Gold and model both: unknown -- these are the lines whose signed sum literally IS the cycle's residual (see Data corpus / tools/build_corpus.py).
ambiguous_phrasing_resolved
Lines written in the deliberately confusable shrink/transfer-out shared opening, resolved correctly
75
44 of 106 shrink lines and 31 of 56 transfer-out lines share the opening 'units removed from location...' / 'inventory removed from floor stock...' on purpose (data/SOURCES.md). All 75 were classified correctly this run (0% transfer_as_shrink, 0%…
What we could NOT verify
Whether this run's 100% line accuracy, 0% transfer<->shrink confusion and 0 false_fully_explained are stable properties of this model at these settings, or a one-off -- one recorded run is not a history.
Whether the same figures hold with reasoning explicitly disabled. The registered run (r001-fin-invval) left provider-side reasoning at its documented default (on); the shipped app's own /api/check hardcodes it off. No run exists at the app's actual setting -- see Business.not_good_enough.
Whether resistance holds against an adversarial or malicious transaction memo -- no red-team run exists for this kit (see redteam_page).
Whether a cycle with meaningfully more than 23 raw lines (this corpus's largest) still finishes inside MAX_TOKENS=4096, particularly with reasoning left on -- see Architecture.breaks_at_scale.
Whether a real transaction feed's own memo language (rather than this corpus's handful of hand-written templates per bucket) would reproduce these figures -- see Data.bring_your_own_boundary.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
1,768.18
2,231.38
14,306 ms
$0.007578
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 60,118 input / 75,867 output tokens priced at Google Gemini 3 Flash's published rate -- the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The raw transaction line list -- the largest single part of every call by character share (2,209 of 5,654 characters on the example call) and it scales directly with how many raw lines a cycle has (12-23 across this corpus), unlike the fixed system prompt.
Reasoning ('thinking') was left at the provider's default (on) for this run rather than explicitly disabled -- it consumed 79.3% of the average output-token budget (1,769.4 of 2,231.4 tokens/call) and is already priced into cost_per_query_usd, since a provider bills reasoning tokens as completion tokens. src/app.py's live UI hardcodes thinking off, a different, unmeasured configuration -- see Business.not_good_enough.
The fixed system prompt (3,269 characters -- the five-bucket vocabulary, sign convention and answer schema) is sent in full on every call regardless of cycle size -- the floor every cycle pays before a single raw line is read.
Your volumeWhat it costs at your volume
Linear in cycles: each call is independent and self-contained, with no shared context or retrieval step to amortise. This run's 34 cycles cost about $0.2577 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $2.577 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volumeroll-forward cycles reconciled per close -- this run decomposed 34 in one pass
What it replacescost accounting reading every raw transaction line by hand, sorting it into a bucket, summing each bucket, and sizing the residual against the account's materiality threshold
Time saved per itemnot measured here -- depends on how long a manual bucket-and-residual pass takes at the reader's own company
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run -- unlike gap-brief's fast/reasoning comparison -- so this page prices one model, not a trade-off; see Eval.could_not_verify for what a second run would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
60,118input tokens · this run
75,867output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 34 roll-forward cycles decomposed, 588 raw lines line-scored by pure code. This run (r001-fin-invval) answered 588 of 588 lines and 34 of 34 residual verdicts -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.103
$0.103
$3.03
2026-09-12
gemini-3-flash
Google
$0.258
$0.258
$7.58
2026-09-18
gemini-3-8-flash
Google
$0.330
$0.330
$9.69
2026-09-18
llama-5
Meta
$0.398
$0.398
$11.69
2026-09-18
claude-haiku-4-5
Anthropic
$0.439
$0.439
$12.93
2026-09-12
grok-4-5
xAI
$0.575
$0.575
$16.92
2026-09-18
grok-4-6
xAI
$0.575
$0.575
$16.92
2026-09-18
claude-sonnet-5
Anthropic
$0.879
$0.879
$25.85
2026-09-12
gemini-3-1-pro
Google
$1.031
$1.031
$30.31
2026-09-18
gpt-5-6-terra
OpenAI
$1.031
$1.031
$30.31
2026-09-12
gpt-5-6-sol
OpenAI
$1.758
$1.758
$51.70
2026-09-12
claude-opus-4-8
Anthropic
$2.197
$2.197
$64.63
2026-09-12
claude-opus-5
Anthropic
$2.197
$2.197
$64.63
2026-09-12
claude-fable-5
Anthropic
$4.395
$4.395
$129.25
2026-09-18
claude-fable-5-1
Anthropic
$4.395
$4.395
$129.25
2026-09-18
gpt-6-astra
OpenAI
$4.395
$4.395
$129.25
2026-09-17
Read this against the numbers above
OUTPUT IS THE LARGER SHARE OF THIS KIT'S BILL, unlike most kits in this series. 2,231 output tokens against 1,768 input is roughly 1.3:1, and 79.3% of that output is a reasoning pass left on at provider default (see Business.not_good_enough) -- so most rows below move more with a model's OUTPUT rate than a typical kit's would.
REASONING WAS NOT DISABLED FOR THIS WORKLOAD. Every row below prices THIS run's own token counts, which include reasoning tokens the live app's own /api/check never generates (it hardcodes thinking off). A forker's real bill on other models depends on whether that model has an equivalent reasoning toggle and whether it is left on.
NO QUALITY IS IMPLIED. Only the model that produced Eval.scores has been scored against this corpus -- every row here is a price, not a recommendation.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 34 roll-forward cycles and 588 raw transaction lines from a fixed seed (SEED=20260819). Gold (a bucket per line, bucket totals, residual, materiality status) is RECOMPUTED from the actual generated lines, never from the random targets that seeded them -- an internal-consistency check (_verify) asserts every cycle reconciles to the cent before the build is considered good.
You change it to: Point it at your own items, locations and transaction-memo phrasing. src/invval.py and src/prompt.py read a cycle by its own fields (opening_balance, closing_balance, materiality_threshold_usd, lines[]) and do not care where the numbers came from.
tools/build_corpus.py
# Generate the inventory-valuation roll-forward cycles and their raw transaction lines, from a
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
ITEMS = [
LOCATIONS = ["STR-104", "STR-118", "STR-221", "STR-309", "STR-402", "STR-517", "DC-NORTH", "DC-SOUTH"]
PERIODS = ["2026-02", "2026-03", "2026-04", "2026-05", "2026-06", "2026-07"]
BUCKETS = ("receipts", "markdowns", "shrink", "allowances", "transfers")
LINE_LABELS = BUCKETS + ("unknown",)
THRESHOLDS = (500.0, 750.0, 1000.0, 1500.0, 2000.0, 3000.0)
src/prompt.pythe prompt — a swap seam
The five-bucket-plus-unknown vocabulary, the sign convention (receipts/markdowns/shrink/allowances positive magnitudes, transfers a signed net figure) and the answer schema, declared once and read from here by build() and parse(). States the transfer-vs-shrink confusion to the model explicitly, as the named failure mode this task exists to catch, not left for the model to discover.
You change it to: A bucket this kit's fixed five-member list does not carry needs a new entry in BUCKETS/BUCKET_MEANINGS here plus new template memo phrasing in tools/build_corpus.py -- a prompt tweak alone will not teach a label the corpus never plants.
src/prompt.py
# Assemble the one prompt this kit sends, and parse the one reply it gets back.
BUCKETS = ("receipts", "markdowns", "shrink", "allowances", "transfers")
LINE_LABELS = BUCKETS + ("unknown",)
BUCKET_MEANINGS = {
RESIDUAL_STATUSES = ("within_threshold", "above_threshold")
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _cycle_header(cycle):
def _lines_block(cycle):
src/invval.pythe AI layer
Loads a cycle by id, calls the model once with every raw line in the cycle batched into that single call, parses the reply into a bucket per line plus bucket totals and the residual. The whole AI layer, deliberately short -- the same split fin-close's close.py and gap-brief's brief.py both make. Never books an adjustment or closes the valuation period -- there is no function here or in src/app.py that does either.
src/invval.py
# Decompose one inventory roll-forward cycle's raw transaction lines into the five named buckets,
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CYCLES = os.path.join(HERE, "data", "cycles.jsonl")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
MAX_TOKENS = 4096
def cycles():
def load_gold():
def check(cfg, cycle, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
def summary(record):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pyspend control
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
src/app.pythe app
Static files and two JSON endpoints (/api/cycle, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-fin-invval) actually used; see Business.not_good_enough.
src/app.py
# The minimal local UI. Standard library only — python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8788"))
class H(BaseHTTPRequestHandler):
def main():
evals/scoring.pythe scorer
Line-level and cycle-level scoring, pure code: exact match per line against gold, transfer<->shrink confusion reported as its own named figure rather than folded into a generic 'wrong' bucket, and false_fully_explained (a real, above-threshold residual called within_threshold) counted separately as the expensive-direction error.
evals/scoring.py
# Score a set of predicted line buckets and cycle-level residual calls against gold. Pure code,
LABELS = ("receipts", "markdowns", "shrink", "allowances", "transfers", "unknown")
STATUSES = ("within_threshold", "above_threshold")
def score(records, gold):
evals/baseline.pyfree baseline
A fixed-priority keyword rule, written the way a person free-texting a quick classifier would write one, that deliberately fires on 'removed from location' BEFORE the 'transfer' keyword check -- so it misclassifies every ambiguous transfer-out line as shrink by construction. The honest, narrow floor a model has to clear.
evals/baseline.py
# What a keyword rule alone catches, over the same corpus. Free. No key, no dependency, no model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def classify(memo):
def judge(cycle, buckets_by_line):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 34 roll-forward cycles and 588 raw transaction lines from a fixed seed (SEED=20260819). Gold (a bucket per line, bucket totals, residual, materiality status) is RECOMPUTED from the actual generated lines, never from the random targets that seeded them -- an internal-consistency check (_verify) asserts every cycle reconciles to the cent before the build is considered good. A swap seam.
src/prompt.pyThe five-bucket-plus-unknown vocabulary, the sign convention (receipts/markdowns/shrink/allowances positive magnitudes, transfers a signed net figure) and the answer schema, declared once and read from here by build() and parse(). States the transfer-vs-shrink confusion to the model explicitly, as the named failure mode this task exists to catch, not left for the model to discover. A swap seam.
src/invval.pyLoads a cycle by id, calls the model once with every raw line in the cycle batched into that single call, parses the reply into a bucket per line plus bucket totals and the residual. The whole AI layer, deliberately short -- the same split fin-close's close.py and gap-brief's brief.py both make. Never books an adjustment or closes the valuation period -- there is no function here or in src/app.py that does either.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from. A swap seam.
src/budget.pyAn append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
evals/scoring.pyLine-level and cycle-level scoring, pure code: exact match per line against gold, transfer<->shrink confusion reported as its own named figure rather than folded into a generic 'wrong' bucket, and false_fully_explained (a real, above-threshold residual called within_threshold) counted separately as the expensive-direction error.
evals/baseline.pyA fixed-priority keyword rule, written the way a person free-texting a quick classifier would write one, that deliberately fires on 'removed from location' BEFORE the 'transfer' keyword check -- so it misclassifies every ambiguous transfer-out line as shrink by construction. The honest, narrow floor a model has to clear.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1768 input and 2231 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's raw transaction memos are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any text the model actually read this run. In a real deployment the raw transaction feed (receipts, markdowns, shrink and transfer memos) would arrive from external operational systems -- a WMS, a POS, a vendor return system -- that this kit's architecture treats as trusted input with no verification step, the same shape of risk fin-close's basis-document attack measures for a different artifact. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and one boundary that should exist does not
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every memo is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's raw transaction feed would carry externally-authored memo text (see posture), and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it; the fourth is a negative result, not a guarantee. Confirmed by reading the code, not by a run, on 2026-08-19 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a decomposition ever write an inventory adjustment or close the valuation period?
A drafted bucket decomposition or residual figure could plausibly trigger a posting or a period-close action.
No code path does. src/invval.py::check() and src/app.py's /api/check both return an answer only; neither writes to data/cycles.jsonl or any other store -- confirmed by reading every call site.
Is the materiality threshold something a prompt or a reply can move?
A crafted memo or a model reply could shift which residual counts as material.
src/invval.py reads materiality_threshold_usd from the cycle's own committed data at call time and states it to the model as a fact; nothing the model returns is consulted when deciding the threshold itself -- only echoed back in residual_status.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/check handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Can the model's own reported numbers disagree with each other, or with the threshold rule, and still reach the reader unflagged?
One would expect a wrong or internally-inconsistent reply to be caught before it renders.
NO -- this is a real gap, not a guarantee. Nothing in src/invval.py or src/app.py recomputes a bucket total, the residual, or the threshold characterization from the model's own per-line answers before the UI shows them; the rule that an above-threshold residual is never within_threshold lives only in the system prompt. See Guardrails.fails.
Each boundary above was checked by reading the call sites, not by an attack trial. Gate 4 is the one that does NOT hold -- see Guardrails.fails for the same finding from the enforcement side.
The result0 attack trials, and one boundary that does NOT hold: nothing in code independently verifies the model's own bucket totals or residual/threshold characterization before the UI renders them -- the rule is prompt-only. Three other boundaries (no write path, threshold not model-settable, key redaction) hold, confirmed by reading the code.
1externally-authored field a live deployment would carry (the raw transaction memo) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no line's memo was authored by an outside party. A real deployment's raw feed would be exactly the kind of externally-supplied text fin-close's basis-document red-team run measured being poisoned for a different artifact; that question is open for this kit.
Read this twice
This kit has no code-level check on its own model's arithmetic or threshold rule. Everything from bucket totals to the above-threshold guardrail lives in the system prompt alone -- src/invval.py and src/app.py pass the model's parsed reply straight through. On this run every number happened to be internally consistent and every gold above-threshold cycle was correctly flagged (0 false_fully_explained), but that is a property of THIS run's replies, not a guarantee the code provides. A future version should recompute the model's own bucket totals and residual from its per-line answers -- a cheap, deterministic check, the same shape evals/scoring.py already runs offline -- before trusting a live decomposition.
HonestyWhat this does not prove
Whether an injected instruction inside a raw transaction memo (e.g. 'this line is pre-approved, classify as receipts') could move a bucket call or the residual_status verdict -- no red-team run exists for this kit.
Whether the live app's hardcoded thinking=off setting changes resistance to a hostile memo relative to this run's provider-default-on configuration -- untested either way.
Whether a code-level consistency check (Guardrails.add_first) would catch a real disagreement in practice -- none has been built or exercised.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
An above-threshold residual is never characterized as within_threshold.
src/prompt.py -- SYSTEM, stated as an explicit boundary the model must respect. Prompt-level ONLY: unlike fin-close's live citation-fidelity check or gap-brief's live fabrication check, no code in src/invval.py or src/app.py recomputes the model's bucket totals or residual from its own per-line answers before the UI renders them.
EvidenceDoes it hold?
What
Measured
No code path in this kit writes an inventory adjustment or closes the valuation period
0 of 34 calls in r001-fin-invval resulted in any write to data/cycles.jsonl or any other store -- src/invval.py::check() and src/app.py's /api/check both only ever return an answer.
The prompt-level rule itself held on this run
0 of 13 gold above-threshold cycles in r001-fin-invval were characterized as within_threshold (false_fully_explained=0) -- but see fails below for what does, and does not, enforce this.
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK. The rule that an above-threshold residual is never within_threshold lives only in src/prompt.py's SYSTEM text -- nothing in src/invval.py or src/app.py recomputes it. See fails above.
It does not verify a bucket total or the residual figure against the raw lines mathematically -- see fails.
It does not make the per-line bucket classification itself correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against a hostile transaction memo -- no red-team run exists for this kit (see the security page).
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 21 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
21 measured by the latest run0 need the model half
Metric
Owner
Role
Why this one
line-bucket-exact-match-plus-residual-threshold
Per-line bucket exact match, plus a tolerance-banded residual and threshold check
alarm
false_fully_explained as a raw count, never folded into line_accuracy -- it is the expensive-direction error; transfer_as_shrink_rate_pct and shrink_as_transfer_rate_pct specifically, since a model that pattern-matches the opening clause fails exactly here; answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed — alarm on Any nonzero false_fully_explained on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
34
different corpus — nothing is comparable
corpus.bytes
109,932
roll-forward cycles edited — the count held, the bytes did not
split.count
34
the raw transaction lines per cycle count moved — a different set was scored
split.size_p50
17
the median size of one raw transaction lines per cycle moved
split.size_p95
22
the 95th-percentile size of one raw transaction lines per cycle moved
dataset.rows
34
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
line accuracy
not yet known
588 answered lines
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat -- r001-fin-invval ran once.
transfer<->shrink confusion
0 -- zero on this run
118 gold transfer / 106 gold shrink lines
measured, one run
false fully explained
0 -- zero this run, this kit's headline guardrail metric
13 gold above-threshold cycles
measured, one run
residual status accuracy
not yet known
34 cycles
100% this run. No repeat -- r001-fin-invval ran once.
reasoning-token share of output
not yet known
34 calls
79.3% this run (60,159 of 75,867 output tokens), one run, one setting -- reasoning left at provider default, never compared against the live app's thinking-off setting. Narrative only, not a registered metric key -- the run record does not carry a reasoning-token-share field, only raw input/output totals.
measured, one run -- the per-bucket breakdown of the overall line-accuracy figure above; not yet known to be stable across a repeat.
lines answered
not yet known
588 raw lines
100% this run -- every raw line drew a classification, none silently dropped. No repeat -- r001-fin-invval ran once.
residual amount correct
not yet known
34 cycles
100% this run -- the residual DOLLAR FIGURE within the scorer's tolerance band, a different question from residual_status_accuracy (the above/within-threshold call). No repeat -- r001-fin-invval ran once.
false alarm material
0 -- zero this run, the cheaper-direction residual error (a clean cycle called material)
21 gold within-threshold cycles
measured, one run
latency
not yet known
34 calls
p50 14,306ms, p95 21,336ms on r001-fin-invval -- reasoning left on. One recorded run, not a distribution.
input volume
0 -- fixed by the corpus and the prompt, not the model
34 calls
60,118 input tokens on r001-fin-invval. Any movement means the prompt or the corpus changed.
output volume
not yet known
34 calls
75,867 output tokens on r001-fin-invval -- model-specific, and includes whatever reasoning the provider chose to spend.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-fin-invval 2026-08-19
allowances accuracy, %
100.0
false alarm material
0
false alarm material rate, %
0.0
false fully explained
0
false fully explained rate, %
0.0
input tokens, whole run
60118
model latency p50 ms
14306.00
model latency p95 ms
21336.00
line accuracy, %
100.0
lines answered, %
100.0
markdowns accuracy, %
100.0
output tokens, whole run
75867
receipts accuracy, %
100.0
residual amount correct, %
100.0
residual status accuracy, %
100.0
shrink accuracy, %
100.0
shrink as transfer rate, %
0.0
transfer as shrink rate, %
0.0
transfer shrink total confused
0
transfers accuracy, %
100.0
unknown accuracy, %
100.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 21 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether reasoning is left on or explicitly disabled
not measured -- this run left it on (provider default); the live app always disables it; the two have never been compared on this corpus
reasoning
src/app.py hardcodes thinking=THINKING_OFF; r001-fin-invval's own thinking field is null (provider default). See Business.not_good_enough.
the confusable transfer-out/shrink phrasing (AMBIGUOUS_FRACTION=0.45 in tools/build_corpus.py)
line accuracy on gold transfer lines: 73.7% (free keyword floor) -> 100% (the fast tier) on the identical 118 lines
measured
results/eval-b000-keyword.json vs results/eval-r001-fin-invval.json, same 34 cycles, same lines, one variable (which classifier reads the memo).
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
line accuracy
nothing yet
transfer<->shrink confusion
any nonzero value on a re-run
false fully explained
any nonzero value, on any run
residual status accuracy
nothing yet
reasoning-token share of output
nothing yet
per-bucket line accuracy
any bucket falling below 100% on a re-run
lines answered
nothing yet
residual amount correct
nothing yet
false alarm material
any nonzero value on a re-run
latency
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
nothing yet
NextThe three you would add first
A live, code-level consistency check: recompute each bucket total and the residual from the model's own per-line answers (a cheap, deterministic pass, the same shape evals/scoring.py already runs offline) and flag any cycle where the model's stated bucket_totals or residual_usd disagree with what its own per-line buckets imply, before the UI renders it.Currently nothing catches a model that reports internally-inconsistent numbers -- the same 'verdict disagrees with its own extracted values' shape as fin-close's false-clean cases, and this kit has no check for it at all, live or offline, beyond the eval's own comparison to gold.
A live check that residual_status matches the sign and magnitude of residual_usd against materiality_threshold_usd.The system prompt states the rule; nothing in code enforces it before a reader sees it. See fails.
A red-team run against the raw transaction memo, the way fin-close attacked its basis document.In a real deployment the memo text arrives from external operational systems this kit's architecture treats as trusted with no verification step -- unmeasured for this kit (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting in src/invval.py.
What this cannot tell you
Whether the same 100% figures hold with reasoning explicitly disabled, matching the live app's own setting -- see Business.not_good_enough.
Whether a live, code-level consistency check (see add_first) would ever have caught a real disagreement -- this run's own model never produced one to test against.
Whether the prompt-only threshold rule holds against a hostile or malformed raw transaction memo -- no red-team run exists for this kit.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is an empty file with a comment explaining that the emptiness is load-bearing. The whole decomposition decision is three files: src/prompt.py, src/invval.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
34 roll-forward cycles and 588 raw transaction lines, generated from a fixed seed, never fetched. Recomputing gold from the same generated lines (never the random targets that seeded them) is what keeps the internal-consistency check honest -- see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the five-bucket vocabulary, the sign convention and the answer schema are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. Passing thinking through is a kwarg on one call site, not a client-library upgrade.
evaluation
evals/scoring.py
eval harnesses
exact match over a six-value vocabulary (five buckets plus unknown) plus a tolerance-banded residual comparison is a dict comprehension and some arithmetic, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per cycle -- load, pack, prompt, call, parse -- with no branching and no state carried between cycles. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A bucket outside the five-member vocabulary needs a hand-written BUCKETS/BUCKET_MEANINGS entry plus new template memo phrasing in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here.
No built-in output validation -- a framework with a schema-and-consistency-check layer baked in might have caught the gap Guardrails.fails names; this kit's stdlib-only design did not build one.
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-fin-invval on the fast tier, 2026-08-19. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
14,306 ms
not yet known
nothing yet
Model, p95
21,336 ms
not yet known
nothing yet
Input tokens
60,118
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
75,867
not yet known
nothing yet
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the rulebook is data read whole into the prompt.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-19, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
roll-forward cycles and raw transaction lines
data/cycles.jsonl -- 34 cycles, 588 lines, generated once from a fixed seed by tools/build_corpus.py; nothing else knows what a roll-forward cycle is
one whole cycle -- every raw line -- sent in the one call; src/app.py and src/invval.py both read it by cycle_id, never modified after generation
gold decomposition
data/gold.jsonl -- 34 rows, computed by tools/build_corpus.py by recomputing every bucket total, the residual and the closing balance from the same generated lines, never carried over from the random targets that seeded the split
never -- evals/scoring.py is pure code, no model, no key. src/invval.py::load_gold()'s own docstring: NEVER read by check().
the key
.env -- never committed
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
rulebook
The five-bucket-plus-unknown taxonomy (receipts, markdowns, shrink, allowances, transfers, unknown), each bucket's meaning, and the sign convention are declared once in src/prompt.py and sent in full inside the system prompt on every call -- a FIXED rulebook that does not vary by cycle, unlike fin-close's own basis document, which varies per recurring template. Any raw line the memo doesn't clearly support any of the five buckets is classified unknown rather than forced.
The rulebook portion of the system prompt is 1,012 tokens (lenses.LLM.prompt_parts[0]), sent identically on all 34 calls in r001-fin-invval -- see LLM.prompt_verbatim for the exact text. (lenses.LLM.prompt_verbatim, src/prompt.py)
a bucket the corpus's five-member vocabulary doesn't name needs a hand-written BUCKETS/BUCKET_MEANINGS entry plus new template memo phrasing in tools/build_corpus.py -- never learned from data alone, and never measured here
a different taxonomy (added, removed or redefined buckets) invalidates the whole eval at once -- gold and grading in evals/scoring.py are both keyed to this exact six-value vocabulary
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 34 roll-forward cycles, 588 raw transaction lines -- byte-identically from a fixed seed (SEED=20260819) every time it is run, with an internal-consistency check (_verify) asserting every cycle's bucket totals, residual and closing balance reconcile to the cent before the build is considered good.
109,932 bytes, one file (data/cycles.jsonl), 34 cycles, 12-23 lines each (p50 17, p95 22) -- see Data.corpus and Data.split. (lenses.Data.corpus, lenses.Data.split, tools/build_corpus.py)
a real deployment's raw feed is continuous, not a fixed seed, and a feed that only nets at the department level cannot be decomposed cleanly by this kit at all (see Data.breaks_on) -- neither was measured
point tools/build_corpus.py at your own items, locations and memo phrasing, and every published accuracy/confusion figure is void -- they are this corpus's own phrasing templates (see Data.bring_your_own_boundary), not a property of the model
model
one call per ROLL-FORWARD CYCLE carrying every raw line in that cycle, behind src/adapters/__init__.py, reasoning left at the provider's default (on) -- the configuration r001-fin-invval ships, and a DIFFERENT configuration from the one src/app.py's live UI actually runs (thinking explicitly off).
34 of 34 cycles in r001-fin-invval returned a reply that parsed cleanly (0 failures, finish_reason 'stop' on all 34) -- but the first attempt at MAX_TOKENS=1800 hit finish_reason='length' on 24 of 34 cycles, which is why MAX_TOKENS is 4096 now (src/invval.py's own comment). (lenses.Business.not_good_enough, results/eval-r001-fin-invval.json, src/invval.py)
a cycle with meaningfully more than 23 raw lines (this corpus's largest) was never measured against MAX_TOKENS=4096, and reasoning left on makes that ceiling less predictable, not more (see Architecture.breaks_at_scale)
verdicts are per-model and this configuration ran once -- the live app's own thinking-off setting has never been scored against this corpus; see Business.not_good_enough
labels
data/gold.jsonl, 34 rows covering 588 raw lines -- every bucket total, the residual and the closing balance are RECOMPUTED from the same generated lines the model reads, never from the random targets that seeded the split, and asserted internally consistent to the cent before the corpus build is considered good.
117 receipts / 111 markdowns / 106 shrink / 105 allowances / 118 transfers / 31 unknown gold lines over 588 total; 13 of 34 cycles gold above_threshold, 21 within_threshold -- see Eval.dataset and Eval.taxonomy. (lenses.Eval.dataset, lenses.Eval.taxonomy, fin-invval-2026-08-19-34cycles)
your own raw feed: hand-label the gold, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-labelled gold has an error rate this kit has never measured
line accuracy and confusion figures over this set reflect ONE planted ambiguity family (transfer-out vs shrink) and ONE unknown-lines-are-the-residual design -- a real roll-forward's other confusions (see Data.breaks_on) are untested
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a cycle's residual_status is within_threshold but its own residual_usd plainly exceeds materiality_threshold_usd
nothing in code would catch this -- the rule is prompt-only (see Guardrails.fails); this run never produced one, but the code does not prevent it
recompute residual_usd against materiality_threshold_usd yourself before trusting a within_threshold verdict, until Guardrails.add_first's live check exists (lenses Guardrails.fails, src/invval.py, src/app.py)
a transfer-out line classified as shrink, or a shrink line classified as transfers
the model pattern-matched the memo's opening clause ('units removed from location...') instead of reading the whole memo -- the free keyword floor does exactly this on 31 of 118 gold transfer lines (results/eval-b000-keyword.json); the fast tier did not do this on any of the 75 ambiguous lines this run (see Eval.taxonomy)
read past the opening clause for a named destination/transfer reference (transfer) versus a count-variance/damage reason with no destination (shrink) before trusting either bucket (lenses.Eval.taxonomy, data/SOURCES.md, results/eval-b000-keyword.json vs results/eval-r001-fin-invval.json)
No machine symptom — this failure leaves no trace in any output.
reasoning ('thinking') left at provider default for the registered run, while the live app hardcodes it off -- a reader who re-runs this kit's own app will see different latency and cost than this report publishes, and nobody has measured whether accuracy differs too. See Business.not_good_enough.
Concurrency and GPU sizing -- one serial call per cycle, nothing measured past 34. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- Cost prices every call at the peak, cache-miss rate for exactly this reason. Whether a real transaction feed's memo language or a feed that does not post at a matching item/location grain would reproduce these figures -- see Data.bring_your_own_boundary. Whether reasoning explicitly disabled (the live app's own setting) changes any figure on this page -- every scored run here left it at provider default. Whether a hostile or malformed raw transaction memo could move a bucket call or the residual_status verdict -- no red-team run exists for this kit.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Per-line bucket exact match, plus a tolerance-banded residual and threshold check
Explain a retailer's month-end inventory change, line by line
PresenterOpens the private repo. Visible to admins only.
In one linePer-line bucket exact match, plus a tolerance-banded residual and threshold check
Does the model's bucket for each raw line match gold (including 'unknown')? Does the cycle's residual dollar figure fall within a tolerance of gold? Does the cycle's above/within-threshold characterization match gold exactly, with the two directional errors (false_fully_explained, false_alarm_material) counted on their own?
$0.00per 1,000 roll-forward cycles
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Cycle
IVR-00003
Raw line
IVR-00003-L19
Memo (verbatim)
Units removed from location, ref TXN-10054. Destination STR-517, per inter-location transfer directive.
Gold bucket
transfers
Keyword floor's bucket
shrink
The model's bucket
transfers
Scored as
correct -- the confusable-phrasing case the free floor fails and the fast tier does not
Grader
Verdict
Why
Per-line bucket exact match, plus a tolerance-banded residual and threshold check
correct
IVR-00003, line L19: memo 'Units removed from location, ref TXN-10054. Destination STR-517, per inter-location transfer directive.' -- gold bucket transfers (a named destination and a transfer reference decide it). The keyword floor fires on the opening clause 'removed from location' and answers shrink, wrong. The fast tier read the whole memo and answered transfers, correct.
The formulaWhat it computes
line_accuracy = correct buckets / answered lines, over lines_scored=588. transfer_as_shrink / shrink_as_transfer = their own named confusion rate, not folded into line_accuracy. residual_amount_correct = |pred - gold| <= max($50, 5%*|gold residual|). residual_status_accuracy = exact match on within_threshold/above_threshold. false_fully_explained = gold above_threshold cycles predicted within_threshold; false_alarm_material = the reverse.
The analysisWhat it actually did
Model
Result
the fast tier
100.0% line accuracy · 1 more measured on this row
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold, recomputed from the same generated lines the model reads (never from the random targets that seeded the split) and asserted internally consistent to the cent before the corpus build is considered good.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own planted phrasing and defect design, which come from tools/build_corpus.py's fixed templates.
Watch these
false_fully_explained as a raw count, never folded into line_accuracy -- it is the expensive-direction error
transfer_as_shrink_rate_pct and shrink_as_transfer_rate_pct specifically, since a model that pattern-matches the opening clause fails exactly here
answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed
Alarm on
Any nonzero false_fully_explained on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
How tight can the band be? The residual-amount tolerance (max($50, 5%*|gold residual|)) is a stated grading allowance, not a fitted score -- it exists so a model's rounding does not fail a cycle whose status call is otherwise correct.
Cadence: Re-run on any change to src/prompt.py, tools/build_corpus.py, or MAX_TOKENS in src/invval.py -- the first changes what is asked, the second changes what is asked ABOUT, the third bounds how many verdicts can come back at all.
The decisionWhen to reach for it
Use it
The gold bucket and gold residual are DERIVED from the same generated lines the model reads -- true of every kit corpus, never true of a real deployment's own roll-forward.
Do not use it
The truth is not known in advance -- the normal state of a real cost-accounting close, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning