Catch where a retailer's demand, supply and finance plans disagree
Your demand, supply and finance teams each keep their own plan, and by review day the numbers have drifted apart. This app lines them up, lists the gaps that matter, and drafts the meeting brief with a likely cause for each, or says unknown.
PresenterOpens the private repo. Visible to admins only.
For the planning analystsCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
A planning team at an outdoor-gear retailer, getting each region ready for its quarterly plan review.
✕Today's manual process
1Open all three plans side by side in a spreadsheet, one line at a time.
2Work out which gaps matter, and which plan never came in.
3Chase the planning notes and the people who wrote them to learn why each number moved.
4One missed gap reaches the review unexplained, and the quarter is signed off on the wrong number.
Every line compared by a person
✓With the app
1The three plans are lined up for every line, automatically.
2Gaps that matter are listed in dollars and percent, with any missing plan named.
3Each gap gets a likely cause, quoting two lines of the planning notes, or it says unknown.
4The meeting gets a short brief that ties the gaps together. People still decide what to do.
People start from a drafted brief
See it work
One real case: what the app reads, step by step
East Region's first-quarter plans: Rain Shells differ by $44,943 across the three, and no supply plan came in for Hydration Gear.
Catch where a retailer's demand, supply and finance plans disagreeReference appBuilt to be shaped to your process
4
1Three plans, one line demand, supply and finance for Rain Shells, a $44,943 gap.
2A plan never sent no supply plan came in for Hydration Gear.
3The planners' notes two notes on Rain Shells explain the gap.
4The gap gets a cause assumption mismatch, cited to those two notes; the rest stay unknown.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch where a retailer's demand, supply and finance plans disagree
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Three teams maintain three views of the same underlying plan -- demand, supply and financial, or any other three-way split a business tracks -- and by the time a review meeting is called, the numbers have quietly diverged. Finding every material gap by hand means opening all three views side by side, line by line, then chasing down whoever wrote the planning notes to work out why, before anyone can even discuss what to do about it. Someone opening all three plan views side by side, line by line, then chasing down whoever wrote the planning notes to work out why a number moved, before anyone can even discuss what to do about it.
Audience
Whoever runs a plan-reconciliation review meeting, and the analysts who prep the itemized gap list beforehand. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual planning cycles
The corpus is 40 planning cycles, 0.07 MB (jsonl 2). A real three-way plan reconciliation names real SKUs, real committed volumes and real revenue numbers -- exactly the material a business will not let leave the building, on top of being internal planning correspondence nobody signs an MIT licence over. Seven scenarios were planted deliberately -- four traceable causes, one genuinely untraceable gap, one missing view, and clean agreement -- to give the eval something real to tell apart: completeness, correct cause-tagging, and honest 'unknown' all need a corpus that plants failures of each kind on purpose.
The corpus
The 40 planning cyclesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your planning cycles. That is the whole change — there is no database to migrate.
One planning cycle, as the model receives itcycles.jsonl · 1 of 40
An itemized list of every material gap, each with a probable cause traced to two source lines -- or an honest 'unknown' -- and a short narrative brief for the meeting. Decision-free by design: nothing here ranks the three plan views or recommends an action.
And when it cannot
A real gap goes unlisted (dropped -- 9 of 146 this run) or a phantom line is listed under a malformed id (fabricated as a duplicate -- 7 of 146) -- both traced this run to the model's own output-formatting slip on 3 of 40 cycles, not a missed read of the notes. See not_good_enough.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
A gap with two clear, item-labelled notes lines explaining it — the fast tier 100% cause-tag agreement on every traceable gap this run measured, with zero fabricated citations.
A gap the notes log genuinely does not explain — the fast tier It said 'unknown' correctly on all 38 gold-unknown gaps this run measured, rather than reaching for the nearest plausible-sounding line.
Feeding the brief's output straight into a system that keys on item_id — add a schema-validation pass before trusting the id field 2 of 40 cycles returned an item_id with the label appended, which broke an exact-match join even though the cause and citations for those same items may have been reasonable -- this run's scorer never got to check.
A cycle with far more planning-notes chatter than this corpus plants — not measured here src/pack.py's 40-line notes cap was never exercised -- every cycle in this corpus ships 6-11 lines.
At a glanceHow the whole thing runs
94%gap completeness recall pct
3,970 msp50, end to end
$0.33per 1,000 planning cycles · Google Gemini 2.5 Flash-Lite
Run once, for real, on 2026-08-18. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch where a retailer's demand, supply and finance plans disagree14 steps · 4 questions · run once, for real · 2026-08-18
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
This kit's actual job -- reconcile three independently-maintained numeric plan views of the same underlying reality, itemize every material gap with a quantified impact and a probable cause traced to real source lines, and draft the meeting brief -- applies to any three-way plan split: a demand/supply/financial plan is one instance, a budget/forecast/actuals set or a hiring-plan/headcount-budget/actual-headcount set are others. The measured completeness, cause-tag-agreement and fabrication figures are properties of THIS corpus's five templated cause-phrasings, its 12% materiality threshold and its 6-11-line notes logs.Corpus lens →
When is this the wrong choice?
Avoid: Nothing -- this is the case the model handles cleanly. That is the case against the best-fitting scenario (“A gap with two clear, item-labelled notes lines explaining it”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
One planted cause per gap. A real gap can have two compounding causes at once; this corpus always plants exactly one (or none, for the untraceable case). 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether the 2-of-40 item_id-formatting slip and the 1-of-40 malformed-JSON reply are stable properties of this model at this settings, or one-off -- one recorded run per tier is not a history. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the reasoning tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-18 — r001-gap-brief. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py) and serves the whole panel -- material gaps, plan views, notes log -- with no call made (python3 -m src.app). Measured on 2026-08-18 with API_KEY blank: two commands, both exit 0, plain python3, no install step -- requirements.txt is deliberately empty. What it cannot do without a key is re-run the eval or press Draft brief; the recorded run renders from the committed result file either way.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
97.5%rows answered
3,970 msp50, end to end
5,141 msp95
2 minclone to first result
What the clock covers. END-TO-END per planning cycle: one HTTP request carrying every material gap plus the cycle's notes log, and the reply parsed to a per-item cause/citation list plus a narrative. Aligning the three plan views and deciding materiality happen before this clock starts and cost no network at all -- pure code over the cycle's own JSON.
Current processWhat it replaces
Someone opening all three plan views side by side, line by line, then chasing down whoever wrote the planning notes to work out why a number moved, before anyone can even discuss what to do about it.
Where it is not good enough
Every dropped or fabricated item on this run traces to one of two mechanical output-formatting failures, never to a judgement error. In 2 of 40 cycles (GB-0006, GB-0009) the model echoed item_id as "IT-2 (Base Layers)" instead of the bare id the schema specified -- that single formatting slip alone accounts for 14 of the run's 16 completeness misses (7 dropped, 7 fabricated as a mismatched duplicate). In 1 of 40 cycles (GB-0017) the reply was malformed JSON -- an unterminated narrative string with no closing brace, despite finish_reason 'stop' and only 265 of a 2,200-token ceiling used, so this was not a budget cutoff -- which zeroed that cycle's two gaps and its narrative outright. Wherever a reply parsed and used the exact schema (37 of 40 cycles), cause-tag agreement, citation fidelity and narrative faithfulness were all 100%: this run's only weakness is output-format discipline, not reasoning. The reasoning tier does not fix this -- it makes it worse: the identical item_id-formatting slip hits 15 of its 40 cycles (vs 2 of 40 on the fast tier), collapsing its completeness recall to 64.4%, at roughly 3.2x the cost. See Cost.cost_by_model.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
the always-unknown floor scores 100 pct completeness and 29.5 pct cause-tag agreement at $0.00 (it never claims a specific cause); the fast-tier model reaches 93.8 pct completeness and 100 pct cause-tag agreement at $0.000326 a cycle but is not clean either — every miss traces to an item-id formatting slip on 3 of 40 cycles, not a wrong judgement. The reasoning tier costs roughly 3.2x more and breaks that same schema 7.5x more often, collapsing its own completeness to 64.4 pct.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the cause vocabulary
src/rubric.py
A cause this kit's fixed five-member list does not carry needs a new CAUSE_VOCAB entry here plus new template notes lines in tools/build_corpus.py -- a prompt tweak alone will not teach it.
the corpus
tools/build_corpus.py
Point it at your own three plan views and your own line items. src/segment.py, src/pack.py and src/prompt.py read cycles by id and do not care where the numbers came from.
the materiality threshold
src/segment.py
MATERIALITY_PCT is a stated judgement, not fitted to this corpus -- change it and re-run the eval, which costs one call per cycle.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 40 planning cycles, 200 line items across three plan views, a planning-notes log per cycle and the gold gap list, from a fixed seed. Seven scenarios are planted, each with a KNOWN true cause (or none, on purpose) -- see data/SOURCES.md.
the cut
src/segment.py
SEAM -- aligns the three plan views for every line item and decides materiality (12% spread, or a missing view by itself). Pure code, no model. The SAME threshold the generator was sized against, imported by both -- see src/segment.py's own header.
the rubric
src/rubric.py
The five-member cause vocabulary and the three graded axes (completeness, cause-tag agreement, narrative faithfulness), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart.
Pack
src/pack.py
Deterministic assembly of one cycle's material-only gap list and its notes log into the context the model sees, capped to a stated budget. Pure code -- a clean item never reaches the model.
the prompt
src/prompt.py
The cause vocabulary, the per-gap answer schema and the narrative instruction, declared once and read from here by the prompt, the parser and the app.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider, stdlib only.
the AI layer
src/brief.py
Loads one cycle, segments and packs its material gaps, calls the model once, parses the reply. The whole AI layer, deliberately short.
the app
src/app.py
Static files and a handful of JSON endpoints on the standard library. Runs the same citation-fidelity check evals/scoring.py runs, on every live draft, before the UI renders it -- not only after the fact in a committed eval result.
Where it breaks at scale
Not on cycle count -- each cycle is independent and the work is linear. It breaks on NOTES LOG LENGTH: src/pack.py caps at 40 lines, untested at that ceiling (every cycle here ships 6-11), and the model has to read the WHOLE log and correlate item to note itself with no pre-filtering. A cycle with a real deployment's volume of planning chatter was not measured, and is exactly where the item_id-formatting slip seen on 2 of 40 cycles here (see Business.not_good_enough) would compound: a longer answer list is more chances for a malformed id to break the join.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Cycle GB-0001, East Region 2026-Q1: five material gaps shown against their three plan views and the cycle's own planning notes, before any call is made.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page after pressing Draft brief with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. No draft is shown -- this is the honest failure state, not a staged one.failureOpen full size →
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
40planning cycles
0.07 MiBjsonl 2
p50 4chars per material gap
$0.00setup · 0.05s
How it is cutWhat one material gap is
One cycle, cut to its material-only gap list by src/segment.py::material_gaps -- no train/test split, every cycle is drafted once.
SetupWhat the setup figure measured
There is no index. 0.05s is tools/build_corpus.py generating 40 cycles, 200 line items and their notes from a fixed seed, with no clock read and no model called.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own planning cycles
This kit's actual job -- reconcile three independently-maintained numeric plan views of the same underlying reality, itemize every material gap with a quantified impact and a probable cause traced to real source lines, and draft the meeting brief -- applies to any three-way plan split: a demand/supply/financial plan is one instance, a budget/forecast/actuals set or a hiring-plan/headcount-budget/actual-headcount set are others. Point tools/build_corpus.py at your own three views, your own line items and your own notes source; src/segment.py, src/pack.py and src/prompt.py read cycles by id and do not change.
⚠︎ And what stops being true when you do: The measured completeness, cause-tag-agreement and fabrication figures are properties of THIS corpus's five templated cause-phrasings, its 12% materiality threshold and its 6-11-line notes logs. A cause phrased differently than this corpus's own templates, a materiality bar set elsewhere, or a notes log an order of magnitude longer -- none of that has been measured, and the 2-of-40 id-formatting slip (Business.not_good_enough) is evidence the model does not always hold a schema perfectly even on this corpus's own modest scale.
What breaks it
One planted cause per gap. A real gap can have two compounding causes at once; this corpus always plants exactly one (or none, for the untraceable case).
One static cycle per line item, not a rolling history. A real deployment would see the same item recur cycle over cycle, sometimes still gapped from last time; this corpus judges every cycle independently.
Two items in the same cycle sharing a root cause. Every item's scenario is drawn independently in tools/build_corpus.py.
A cause outside the five-member vocabulary. A real reconciliation meeting will eventually name a cause this kit's fixed list does not carry.
A notes log at real deployment volume. This corpus ships 6-11 lines per cycle; src/pack.py's 40-line cap was never exercised.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
2,920
772
gaps
757
200
notes
739
195
Total
1,167
This is the cost lesson as arithmetic: of the 1,167 tokens assembled, 772 are instructions — 66% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged verbatim by evals/run.py on r001-gap-brief -- the literal user message sent for cycle GB-0001, with the fixed system prompt from src/prompt.py prepended.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are drafting the gap brief for a plan-reconciliation review meeting. You are given one planning cycle's MATERIAL gaps -- line items where three independently-maintained plan views (demand, supply, financial) disagree by enough to matter, already identified by code -- and that cycle's own planning notes log.
For EACH gap in the list, decide which of these five causes applies, using ONLY the notes log -- never outside knowledge, never a plausible-sounding guess:
timing_lag one view has not been refreshed since a change (a price move, a promo date, a schedule slip) that the other views already reflect.
assumption_mismatch two views were built on different stated assumptions for the same line -- a different price point, a different promo condition, a different scope of what counts.
data_entry_error a transcription or unit mistake in how a number was carried from one planning system into another.
scope_mismatch one view rolls in a sub-line item (an allowance, an add-on, a phase-out) that another view excludes.
unknown the notes do not support any of the four causes above for this specific line item. This is the correct answer when the evidence is not there -- guessing a specific cause is a worse answer than saying so.
If the notes log supports a specific cause for that item, name it and cite the exact two notes lines (verbatim) that together support it. If the notes log does NOT support a specific cause for that item -- no line mentions it, or the lines that do mention it don't add up to one of the four named causes -- the cause is 'unknown' and the two citation fields must be empty strings. Do not cite a line that does not actually name or clearly concern that item just because it is the closest-sounding one available: an unsupported 'unknown' is the correct, honest answer and is graded as such; a cause with a citation that is not really about that item is graded as a fabrication, which is worse.
Each gap also carries a missing_view field in the input, set to a view name when that view was never submitted this cycle, or null when all three views are present. Echo that exact state back in your answer for that gap -- true if a view is missing, false if not.
After the per-gap entries, write a short narrative (3-6 sentences) for the meeting audience that ties the material gaps together. Every number you state in the narrative -- a dollar figure, a percentage, a count of gaps -- must be a number that is actually present in the gap list you were given; do not compute, round, or restate a figure that changes its value.
NEVER recommend which plan view is correct, and never rank or suggest an action to take. That decision belongs to the humans in the meeting this brief is for -- your job is limited to itemizing, tagging a probable cause (or saying unknown), and describing what the gaps are, never what should be done about them.
Cycle GB-0001 -- East Region, 2026-Q1
MATERIAL GAPS (5)
------------------
- IT-1 (Cookset Accessories): demand=$148289.57, supply=$148677.48, financial=$200256.90 | spread $51967.33 (35.0%) | missing_view=None
- IT-2 (Trail Nutrition): demand=$140308.99, supply=$146057.96, financial=not submitted | spread $5748.97 (4.0%) | missing_view=financial_plan_usd
- IT-3 (Camp Cookware): demand=$159786.21, supply=$104050.13, financial=$161477.80 | spread $57427.67 (35.9%) | missing_view=None
- IT-4 (Rain Shells): demand=$93117.88, supply=$136714.81, financial=$138061.13 | spread $44943.25 (32.9%) | missing_view=None
- IT-5 (Hydration Gear): demand=$176986.26, supply=not submitted, financial=$177722.87 | spread $736.61 (0.4%) | missing_view=supply_plan_usd
PLANNING NOTES LOG
-------------------
- IT ticket open for the planning tool's export button; workaround is CSV download.
- Finance close calendar shifted two days for the holiday -- see the shared calendar invite.
- Reminder: cycle sign-off is due Friday EOD, no exceptions this quarter.
- Flagging Rain Shells -- confirmed the assumption gap on demand plan is a real difference in what each side assumed, not a data issue.
- Regional lead is out next week; route urgent items to the deputy instead.
- Vendor scorecard refresh landed in the shared drive, unrelated to this cycle's numbers.
- Rain Shells: demand side is modeling this at a different price/promo assumption than supply side used -- the two aren't reconciled on inputs yet.
Return a JSON object with two keys: "gaps", a list with exactly 5 entries, one per gap above in the same order, each {"item_id": <id>, "cause": <one of: timing_lag, assumption_mismatch, data_entry_error, scope_mismatch, unknown>, "citation_1": <verbatim notes line, or "">, "citation_2": <verbatim notes line, or "">, "missing_view": <true or false>, "note": <one sentence>}; and "narrative", a string of 3-6 sentences.
Answer with the JSON object only.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"gaps": [
{
"item_id": "IT-1",
"cause": "unknown",
"citation_1": "",
"citation_2": "",
"missing_view": false,
"note": "No notes log entry concerns Cookset Accessories."
},
{
"item_id": "IT-2",
"cause": "unknown",
"citation_1": "",
"citation_2": "",
"missing_view": true,
"note": "Financial plan not submitted; no notes log entry concerns Trail Nutrition."
},
{
"item_id": "IT-3",
"cause": "unknown",
"citation_1": "",
"citation_2": "",
"missing_view": false,
"note": "No notes log entry concerns Camp Cookware."
},
{
"item_id": "IT-4",
"cause": "assumption_mismatch",
"citation_1": "Flagging Rain Shells -- confirmed the assumption gap on demand plan is a real difference in what each side assumed, not a data issue.",
"citation_2": "Rain Shells: demand side is modeling this at a different price/promo assumption than supply side used -- the two aren't reconciled on inputs yet.",
"missing_view": false,
"note": "Two notes explicitly state a confirmed assumption gap between demand and supply for Rain Shells."
},
{
"item_id": "IT-5",
"cause": "unknown",
"citation_1": "",
"citation_2": "",
"missing_view": true,
"note": "Supply plan not submitted; no notes log entry concerns Hydration Gear."
}
],
"narrative": "This cycle shows five material gaps across three product families, with two items missing a plan view. The only clearly identified cause is for Rain Shells, where an assumption mismatch is confirmed between demand and supply on pricing and promo inputs. The other gaps—Cookset Accessories, Trail Nutrition, Camp Cookware, and Hydration Gear—have no supporting notes log evidence and are therefore tagged as unknown. The spread is $51,967.33 for Cookset Accessories, $5,748.97 for Trail Nutrition, $57,427.67 for Camp Cookware, and $44,943.25 for Rain Shells, with Hydration Gear having a spread of only $736.61. Given the lack of notes on the remaining items, no additional causal determination is possible without further input."
}
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch where a retailer's demand, supply and finance plans disagree — 40 planning cycles. Two tiers of one model family answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: the item set the brief covers is compared to the gold material-gap set (completeness), each matched item's cause is compared to gold (cause-tag agreement, split by whether gold's own cause was traceable or unknown), each non-unknown cause's two citations are checked as real, item-relevant substrings of that cycle's own notes (fabrication), and every quantified claim in the narrative is checked against the packed gap numbers (faithfulness). No judge model and no rubric scored by a person. The same module scores the app's live citation check.
40planning cycles
40source documents
2model tiers
80graded answers
1grading method
MeasurementsWhat was measured
COUNTED137 · 94 / 146gap completeness recall pct — material gaps in the gold setDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED137 · 94 / 144gap completeness precision pct — gaps the brief itemizedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED137 · 94 / 137cause tag agreement pct — gaps the brief correctly itemizedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 / 137fabricated cause — gaps the brief correctly itemizedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED39 · 40 / 39narrative faithfulness pct — cycles that produced a narrativeDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The method is pure comparison and cannot be wrong about itself; the risk is in the labels. Every gold cause and citation pair is decided by tools/build_corpus.py at generation time, before src/segment.py's materiality check ever runs on the same cycle -- so a model's cause-tag agreement is a measured question, not a tautology.
521.6output tokens · the fast tier · 3,970 ms p50
544.1output tokens · the reasoning tier · 6,532 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.6× as long, and lands one row apart on 40. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One planning cycle
1,000 planning cycles
Share that is the prompt
Google Gemini 2.5 Flash-Lite the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.10 / $0.40
$0.000325
$0.33
36%
Amazon Nova Pro on Bedrock a frontier tier, for the other end of the range
$0.80 / $3.20
$0.002604
$2.60
36%
Same work, 8× the bill
The same planning cycles, the same tokens — only the rate card changed. And on either card about 36% of what you pay is the prompt this pipeline sends, not the answer it writes.
Rates checked 2026-08-18. The provider that actually ran both scored evals is kept out of these tables per this estate's naming rule, so nothing here is what was actually paid -- the real spend for this kit's build (well under $0.05 across both tiers and the free baseline) is recorded in the commit history, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing -- an item set is compared by code, a cause string is compared by code, and a citation is a substring check. A measured $0.00, not an unpriced one.
The gradersOne way to grade, and why it is the only one
The floor gets completeness and missing-view echo for free -- both are pure copying of what it was handed -- and its narrative never invents a number, so it is faithful by construction. What it structurally cannot do is tell a traceable cause from an untraceable one: it always says unknown, so its cause-tag agreement (29.5%) is exactly the gold set's own unknown share -- the fast tier's 100% is the real, measured gap a model closes.
the fast tier 93.8% gap completeness recall · the reasoning tier 64.4% gap completeness recall · 3 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Partial, and mostly about format, not judgement. Both tiers are expected to reach 100% cause-tag agreement on whatever they correctly itemize, because the task per gap is small and the notes are short -- so the labelled set's real power is in catching format discipline (the id-echo slip, the malformed-JSON cycle), which one run each cannot yet say is a stable property of either tier or a one-off. See Guardrails.bands.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
A gap with two clear, item-labelled notes lines explaining it
the fast tier
100% cause-tag agreement on every traceable gap this run measured, with zero fabricated citations.
Nothing -- this is the case the model handles cleanly.
A gap the notes log genuinely does not explain
the fast tier
It said 'unknown' correctly on all 38 gold-unknown gaps this run measured, rather than reaching for the nearest plausible-sounding line.
Assuming a model will always resist guessing -- this corpus's notes are short (6-11 lines); a longer log was not measured.
Feeding the brief's output straight into a system that keys on item_id
add a schema-validation pass before trusting the id field
2 of 40 cycles returned an item_id with the label appended, which broke an exact-match join even though the cause and citations for those same items may have been reasonable -- this run's scorer never got to check.
Trusting item_id as a clean key with no validation, on any run.
A cycle with far more planning-notes chatter than this corpus plants
not measured here
src/pack.py's 40-line notes cap was never exercised -- every cycle in this corpus ships 6-11 lines.
Assuming completeness and cause-tag agreement hold at a real deployment's notes volume without re-measuring.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
correct_traceable_cause
Correctly itemized with the right traceable cause
99
GB-0001, IT-4 (Rain Shells): gold cause assumption_mismatch, model cause assumption_mismatch, both citations real and item-relevant -- see Eval.example_row.
correct_unknown
Correctly said 'unknown' rather than guessing
38
GB-0001, IT-1 (Cookset Accessories): a genuine 35.0% spread with no notes-log line about it. Model: "No notes log entry concerns Cookset Accessories" -- cause unknown, no citations offered.
id_format_mismatch
Item dropped or duplicated because the model appended the label to the id
14
GB-0006 and GB-0009: the model returned item_id "IT-2 (Base Layers)" instead of the schema's bare "IT-2" for every gap in both cycles -- an exact-match join correctly counts these as both a dropped gap (the real id never appears) and a fabricated one (the…
parse_failure
Whole-cycle JSON parse failure
2
GB-0017: the reply's narrative string was never closed with a quote before the final brace, despite finish_reason 'stop' and only 265 of 2,200 output tokens used -- not a budget cutoff. Both of that cycle's material gaps score as dropped, and no narrative was…
fabricated_cause
A non-unknown cause with a citation that was not real and relevant
0
None on either scored tier. Every non-unknown cause cited two real, item-relevant lines from that cycle's own notes.
What we could NOT verify
Whether the 2-of-40 item_id-formatting slip and the 1-of-40 malformed-JSON reply are stable properties of this model at this settings, or one-off -- one recorded run per tier is not a history.
Cause-tag agreement and citation fidelity at a notes-log volume larger than this corpus's 6-11 lines per cycle.
Whether a stricter prompt instruction about the exact item_id format would eliminate the formatting slip, or just move it somewhere else -- not tried here.
Two compounding causes on one gap, or two gaps sharing a root cause -- this corpus never plants either.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 2.5 Flash-Lite
Amazon Nova Pro on Bedrock
the fast tier
1,168.6
521.6
3,970 ms
$0.000325
$0.002604
the reasoning tier
1,168.6
544.1
6,532 ms
$0.000334
$0.002676
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. Figures above are projected onto the cheapest published card, not the real spend.
Cost driversWhat actually moves the bill
The system prompt, the single largest part at 2920 of about 4416 characters (~66%) -- the fixed cause vocabulary and instructions cost more per call than any one cycle's own gap list or notes.
Notes-log length: every line in the cycle's own notes is sent whether or not it explains any gap, by design (src/pack.py never pre-filters).
Output length scales with how many material gaps a cycle has (2-5 here, ~130-230 output tokens per gap) -- a busier reconciliation cycle costs more per call, linearly.
Your volumeWhat it costs at your volume
Linear in cycles and cheap on the fast tier: 40 cycles cost about $0.0124 projected onto this card, so ten times the set is about $0.124. Each call is independent and self-contained -- no shared context or retrieval step to amortise.
Where pricing changes shape
Your return, with your numbers
Volumeplanning cycles reconciled per review period -- this run drafted 40 in one pass per tier
What it replacesa person opening all three plan views side by side and chasing down the planning notes to explain each gap by hand
Time saved per itemnot measured here -- depends on how long a manual gap reconciliation takes at the reader's own company
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose answer might be 'the free floor already wins part of this', which on completeness and missing-view echo it very nearly does. The reasoning tier was run too, at roughly 3.2x the price, and lost on the one axis that moved (completeness) for no gain on the two that held (cause-tag agreement, narrative faithfulness) -- see cost_by_model.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
46,745input tokens · this run
20,866output tokens
$0.012what it actually cost
Every fast-tier model number on these pages: 40 cycles, one call each.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.034
$0.034
$0.86
2026-09-12
gemini-3-flash
Google
$0.086
$0.086
$2.15
2026-09-18
gemini-3-8-flash
Google
$0.113
$0.113
$2.83
2026-09-18
llama-5
Meta
$0.147
$0.147
$3.68
2026-09-18
claude-haiku-4-5
Anthropic
$0.151
$0.151
$3.78
2026-09-12
grok-4-5
xAI
$0.219
$0.219
$5.47
2026-09-18
grok-4-6
xAI
$0.219
$0.219
$5.47
2026-09-18
claude-sonnet-5
Anthropic
$0.302
$0.302
$7.55
2026-09-12
gemini-3-1-pro
Google
$0.344
$0.344
$8.60
2026-09-18
gpt-5-6-terra
OpenAI
$0.344
$0.344
$8.60
2026-09-12
gpt-5-6-sol
OpenAI
$0.604
$0.604
$15.11
2026-09-12
claude-opus-4-8
Anthropic
$0.755
$0.755
$18.88
2026-09-12
claude-opus-5
Anthropic
$0.755
$0.755
$18.88
2026-09-12
claude-fable-5
Anthropic
$1.511
$1.511
$37.77
2026-09-18
claude-fable-5-1
Anthropic
$1.511
$1.511
$37.77
2026-09-18
gpt-6-astra
OpenAI
$1.511
$1.511
$37.77
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate.
Rates are as of each row's own rates_as_of date and nothing checks them against the vendors on an ongoing basis.
Token counts are not comparable across providers -- a sibling kit measured two tiers of one vendor counting a byte-identical prompt differently.
Nothing here includes retries; this kit's two real runs had very few (finish_reason was 'stop' on 79 of 80 calls; one fast-tier reply was malformed JSON, not a retry).
A cheaper row is not a cheaper kit -- the free always-unknown floor costs $0.00 and its completeness figures are published beside every model row in Eval.baseline.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Four of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 40 planning cycles, 200 line items across three plan views, a planning-notes log per cycle and the gold gap list, from a fixed seed. Seven scenarios are planted, each with a KNOWN true cause (or none, on purpose) -- see data/SOURCES.md.
You change it to: Point it at your own three plan views and your own line items. src/segment.py, src/pack.py and src/prompt.py read cycles by id and do not care where the numbers came from.
tools/build_corpus.py
# Generate the planning cycles, their three plan views, their planning notes, and the gold gap
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260818
CATEGORIES = [
BUSINESS_UNITS = ["North Region", "South Region", "East Region", "West Region", "Direct-to-Consumer"]
PERIODS = ["2026-Q1", "2026-Q2", "2026-Q3", "2026-Q4"]
SCENARIOS = ["clean", "timing_lag", "assumption_mismatch", "data_entry_error", "scope_mismatch",
WEIGHTS = [0.28, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12]
NOISE_LINES = [
src/segment.pythe cut — a swap seam
SEAM -- aligns the three plan views for every line item and decides materiality (12% spread, or a missing view by itself). Pure code, no model. The SAME threshold the generator was sized against, imported by both -- see src/segment.py's own header.
You change it to: MATERIALITY_PCT is a stated judgement, not fitted to this corpus -- change it and re-run the eval, which costs one call per cycle.
src/segment.py
# SEAM 1 -- the cut. Aligning three independently-maintained plan views for one line item and
VIEWS = ("demand_plan_usd", "supply_plan_usd", "financial_plan_usd")
MATERIALITY_PCT = 12.0
def align(item):
def align_cycle(cycle):
def material_gaps(cycle):
src/rubric.pythe rubric — a swap seam
The five-member cause vocabulary and the three graded axes (completeness, cause-tag agreement, narrative faithfulness), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart.
You change it to: A cause this kit's fixed five-member list does not carry needs a new CAUSE_VOCAB entry here plus new template notes lines in tools/build_corpus.py -- a prompt tweak alone will not teach it.
src/rubric.py
# The rubric this kit's brief is graded against -- the sections a gap brief must carry, the fixed
CAUSE_VOCAB = (
CAUSE_MEANINGS = {
RUBRIC_AXES = (
FABRICATION_GUARDRAIL = (
src/pack.pyPack
Deterministic assembly of one cycle's material-only gap list and its notes log into the context the model sees, capped to a stated budget. Pure code -- a clean item never reaches the model.
src/pack.py
# SEAM 2 -- Pack. Deterministic assembly of one cycle's material gaps and its notes log into the
MAX_NOTES = 40
MAX_GAPS = 20
def pack(cycle, notes, gaps):
src/prompt.pythe prompt
The cause vocabulary, the per-gap answer schema and the narrative instruction, declared once and read from here by the prompt, the parser and the app.
src/prompt.py
# Assemble the one prompt this kit sends per cycle, and parse the one reply it gets back.
MAX_TOKENS = 2200
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _gap_block(packed):
def _notes_block(packed):
def build(packed, prompt=DEFAULT_PROMPT):
def parse(raw):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider, stdlib only.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/brief.pythe AI layer
Loads one cycle, segments and packs its material gaps, calls the model once, parses the reply. The whole AI layer, deliberately short.
src/brief.py
# SEAM 3 -- the AI layer. Loads one cycle, segments and packs its material gaps, calls the model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
def _read_jsonl(path):
def cycles():
def cycles_by_id():
def notes_by_id():
def gold_by_id():
def draft(cfg, cycle, notes, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
src/app.pythe app
Static files and a handful of JSON endpoints on the standard library. Runs the same citation-fidelity check evals/scoring.py runs, on every live draft, before the UI renders it -- not only after the fact in a committed eval result.
src/app.py
# The minimal local UI. Standard library only -- python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8783"))
def _corpus():
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 40 planning cycles, 200 line items across three plan views, a planning-notes log per cycle and the gold gap list, from a fixed seed. Seven scenarios are planted, each with a KNOWN true cause (or none, on purpose) -- see data/SOURCES.md. A swap seam.
src/segment.pySEAM -- aligns the three plan views for every line item and decides materiality (12% spread, or a missing view by itself). Pure code, no model. The SAME threshold the generator was sized against, imported by both -- see src/segment.py's own header. A swap seam.
src/rubric.pyThe five-member cause vocabulary and the three graded axes (completeness, cause-tag agreement, narrative faithfulness), declared once and imported by the prompt, the app and the scorer so they can never silently drift apart. A swap seam.
src/pack.pyDeterministic assembly of one cycle's material-only gap list and its notes log into the context the model sees, capped to a stated budget. Pure code -- a clean item never reaches the model.
src/prompt.pyThe cause vocabulary, the per-gap answer schema and the narrative instruction, declared once and read from here by the prompt, the parser and the app.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider, stdlib only. A swap seam.
src/brief.pyLoads one cycle, segments and packs its material gaps, calls the model once, parses the reply. The whole AI layer, deliberately short.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1168 input and 521 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface. Every fact in the prompt -- the three plan views, the notes log, the materiality decision -- is computed directly from this kit's own generated corpus (tools/build_corpus.py, src/segment.py); nothing arrives from outside the codebase the way an inbox message or an ingested document does on sibling kits.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. src/app.py strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- there is no surface to attack
An indirect prompt injection needs a field an outside party controls that reaches the prompt. This kit has none, so the four rows below are boundaries confirmed by reading the code, not payloads run through it. Confirmed by reading the code, not by a run, on 2026-08-18 -- no attack run was fired; see the experiment section for why.
Boundary checked
What could go wrong
What the code guarantees
Does the brief's cause tag or citation ever reach a plan-write store?
A drafted cause could plausibly trigger a write back to one of the three plan values, or a 'corrected' number.
No code path does. src/app.py's /api/draft and evals/run.py both return an answer only; neither writes to data/cycles.jsonl or any other store -- confirmed by reading every call site (see Guardrails.holds).
Can a non-'unknown' cause reach a reader with a fabricated or off-item citation?
A model could assert a specific cause and a plausible-looking but fabricated or off-item citation, which a reader would trust as evidence.
src/app.py runs the identical fabrication check evals/scoring.py grades the real run with (citation_is_real and citation_is_relevant) on every live draft, before the UI renders it, and marks any non-unknown cause whose citations fail either check -- confirmed by reading /api/draft.
Is the materiality/itemization decision something a prompt or a reply can move?
A crafted gap description could shift which spread counts as material, widening or narrowing what gets itemized.
src/segment.py::MATERIALITY_PCT is a module-level constant, read once at import time -- nothing the model returns is consulted when deciding which gaps are material; the packed gap list is built before the model is ever called.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/draft handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Each boundary above was checked by reading the call sites, not by an attack trial -- there is no untrusted field to construct a payload against (see posture).
The result0 attack trials, by design -- but 4 boundaries around the model's reply hold, confirmed by reading the code, and the fabrication guardrail runs live, not only in a committed eval result.
0untrusted input fields identified
0 of 0attack trials run
n/adecision-flip resistance -- no such surface exists
This kit has no field an outside party controls that reaches the prompt -- the three plan views and the notes log are all computed from data tools/build_corpus.py generated, never from correspondence, a document, or any other externally supplied text. A future version that ingests real planning notes from a live external system would reopen this question and should be attacked before shipping.
Read this twice
This is a code read, not an attack run. This kit has no ingested document and no external correspondence -- every fact in the prompt is computed from its own generated corpus, so there is nothing for an outside party to write an instruction into. What the four boundaries above guarantee is narrower: even a wrong or adversarial model REPLY could not write a plan value, could not move the materiality threshold, and could not show a reader a fabricated citation without it being flagged, because none of those paths trusts the model's output unchecked.
HonestyWhat this does not prove
Whether a future version of this kit that ingests real planning-notes text from a live external system (rather than a generated corpus) would reopen an injection surface. Not applicable to the shipped version; worth re-examining if this kit is ever pointed at a live notes feed.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
A non-'unknown' cause is shown only alongside its two citations, and every citation is checked as a real, item-relevant line from that cycle's own notes before the UI renders it. No code path in this kit writes to a plan value or ranks the three views.
src/app.py's /api/draft endpoint and evals/run.py both return a per-item cause/citation/note plus a narrative; neither writes to data/cycles.jsonl or any other store. There is no plan-write or apply path in this kit at all.
EvidenceDoes it hold?
What
Measured
No code path in this kit writes to a plan value
0 of 80 calls across both scored tiers (r001-gap-brief, r002-gap-brief-pro) resulted in any write to data/cycles.jsonl or any other store -- src/app.py's /api/draft and evals/run.py both only ever return an answer.
Every non-unknown cause's citations are checked before the reader sees them
src/app.py runs the fabrication check on every live draft; across both scored tiers, 0 of 231 matched, non-unknown-cause gaps (137 fast-tier + 94 reasoning-tier) had a citation that failed it -- the guardrail had nothing to catch on this run, but it ran on all of them.
The limitWhat a guardrail is not
IT IS NOT A RATE LIMIT, A POLICY CHECK, OR A GUARD THAT WITHHOLDS A REPLY FROM BEING SHOWN. It flags a fabricated citation in the UI; it does not suppress the brief -- a reader still sees the flagged line, marked, per the UI's own '.cite.bad' treatment.
It does not make the cause-tag judgement itself correct -- see the id_format_mismatch and parse_failure findings in Eval.taxonomy, which this guardrail does nothing to prevent.
It does not validate the narrative's numbers live, independent of the citation check -- narrative faithfulness is graded by evals/scoring.py after the fact in the eval harness, not enforced on a live draft the way the citation check is. See add_first.
A future version wired to a real plan-write system would need a real, separate approval gate before writing anything; see add_first.
WatchedWhat is watched, and why that one
2runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 24 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
15 measured by the latest run9 need the model half
Metric
Owner
Role
Why this one
cause-and-citation-against-gold
The brief's per-item cause tag and citations against that cycle's gold gap list, and its narrative against the packed gap numbers
alarm
dropped_gap; fabricated_gap; wrong_cause; fabricated_cause — alarm on Any fabricated_cause on any run -- the guardrail this kit's own atlas facet sheet names first. Both scored tiers had zero.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
40
different corpus — nothing is comparable
corpus.bytes
75,498
planning cycles edited — the count held, the bytes did not
split.count
40
the material gaps count moved — a different set was scored
split.size_p50
4
the median size of one material gap moved
split.size_p95
5
the 95th-percentile size of one material gap moved
dataset.rows
40
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.05
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cycles 40, gold_material_total 146, thinking False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
gap completeness
not yet known
146 material gaps
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat -- the fast and reasoning tiers are different models, not two runs of one.
cause-tag agreement
0 -- both tiers hit 100% on the items they matched (traceable and unknown gold causes alike), despite being different models
137 (fast) / 94 (reasoning) matched items
measured, both runs
fabricated cause
0 -- zero on both tiers, this run's headline guardrail metric
137 / 94 matched items
measured, both runs
missing-view echo
0 -- 100% on both tiers, a pure input-to-output copy
18 (fast) / 13 (reasoning) items with a missing view
measured, both runs -- every item whose gold data flagged a missing view had that flag echoed back correctly
narrative faithfulness
0 -- 100% on both tiers wherever a narrative was produced
39 (fast) / 40 (reasoning) cycles with a narrative
measured, both runs -- the one fast-tier cycle with no narrative (GB-0017, the parse failure) is excluded from this rate, not counted as unfaithful
item_id-format-mismatch rate
not yet known
40 cycles
2 of 40 cycles (fast tier) vs 15 of 40 (reasoning tier), one run each, not repeated.
input volume
0 -- identical across models by construction
40 calls
46,745 input tokens on BOTH scored runs, to the token, because prompt assembly is pure code and model-independent.
output volume
not yet known
40 calls
20,866 output tokens on the fast tier, 21,763 on the reasoning tier -- close, but model-specific unlike input.
latency
not yet known
40 calls
p50 3,970ms then 6,532ms, p95 5,141ms then 9,568ms -- the reasoning tier is slower for measurably WORSE completeness. One recorded run per tier, not a distribution.
HistoryRun history
2 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-gap-brief 2026-08-18
r002-gap-brief-pro 2026-08-18
cause tag agreement, %
100.0
100.0
cause tag agreement traceable, %
100.0
100.0
cause tag agreement unknown, %
100.0
100.0
fabricated cause
0
0
false negative
9
52
false positive
7
52
gap completeness precision, %
95.1
64.4
gap completeness recall, %
93.8
64.4
input tokens, whole run
46745
46745
model latency p50 ms
3970.00
6532.00
model latency p95 ms
5141.00
9568.00
missing view echo, %
100.0
100.0
narrative faithfulness, %
100.0
100.0
output tokens, whole run
20866
21763
true positive
137
94
not a time series No two of these 2 runs measured the same system — they differ on the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 2 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
which model drafts the brief
gap-completeness recall 93.8% -> 64.4%, cost roughly 3.2x higher, latency roughly 1.6x higher -- every headline number gets worse or costlier on the pricier tier except cause-tag agreement and narrative faithfulness, which hold at 100% on both
measured
r001-gap-brief (fast) vs r002-gap-brief-pro (reasoning), both with thinking disabled, same corpus, same prompt -- one variable, and it moved completeness sharply in the wrong direction.
whether the reply complies with the exact item_id schema
a cycle's true-positive count from matching its material-gap count down to zero -- a single formatting choice, not a judgement difference, decides whether a cycle's gaps count at all
measured
per_cycle breakdown in both result files: every cycle with false_positive == false_negative and true_positive == 0 is this exact failure mode, on 2 of 40 (fast) and 15 of 40 (reasoning) cycles.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
gap completeness
nothing yet
cause-tag agreement
any value under 100% on a re-run
fabricated cause
any nonzero value, on any run
missing-view echo
any value under 100% on a re-run
narrative faithfulness
any value under 100% on a re-run
item_id-format-mismatch rate
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
nothing yet
latency
nothing yet
NextThe three you would add first
A live narrative-faithfulness check, the same shape as the live citation checkThe citation-fidelity check runs live in src/app.py; the narrative-faithfulness check (evals/scoring.py::narrative_faithfulness) currently only runs in the eval harness, not on a live draft -- a reader of the UI has no equivalent flag today.
A real apply path behind an explicit human-approval stepPer the atlas row's own guardrail ('no plan-write path'), nothing here should ever write a plan value without a person confirming it first -- and there is currently no apply path of any kind to gate.
A schema-validation pass on item_id before any downstream system trusts it2 of 40 fast-tier cycles (15 of 40 reasoning-tier) returned an item_id with the label appended, which this run's own exact-match scorer correctly rejected -- a real integration keying on item_id needs the same rejection before it trusts the field.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/baseline.py (free) on any threshold change in src/segment.py. Re-run evals/run.py (paid, one call per cycle) on any change to src/prompt.py, src/segment.py, src/rubric.py or the corpus generator.
What this cannot tell you
One recorded run per tier is not a history -- whether the id-formatting slip rate is stable across repeated runs on the same corpus has not been measured.
Whether a stricter prompt instruction about the exact item_id format would eliminate the formatting slip on either tier -- not tried here.
Whether a real deployment's own notes-log volume (likely far more than this corpus's 6-11 lines per cycle) would change the fabrication or completeness figures.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
This kit has no framework and no dependencies at all -- json, statistics, re, urllib and http.server from the standard library. requirements.txt is an empty file with a comment explaining that the emptiness is load-bearing.
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the outcome vocabulary and grading
src/rubric.py, evals/scoring.py
eval harnesses (promptfoo, DeepEval)
five causes, three graded axes and a substring check is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per cycle -- align, pack, prompt, call, parse, score -- with no branching and no state carried between cycles. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A cause outside the five-member vocabulary needs a hand-written CAUSE_VOCAB entry plus new template notes lines in tools/build_corpus.py, rather than being configured declaratively.
There is no built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here.
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-gap-brief on the fast tier, 2026-08-18. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
3,970 ms
not yet known
nothing yet
Model, p95
5,141 ms
not yet known
nothing yet
Input tokens
46,745
0 -- identical across models by construction
any change without a corresponding change to prompt or corpus
Output tokens
20,866
not yet known
nothing yet
No movement column. Not one of the 1 earlier run on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-gap-brief3,970 ms
r002-gap-brief-pro6,532 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — one document is one unit, whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-18, across 2 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
planning cycles and plan views
data/cycles.jsonl -- 40 cycles, 200 line items, generated once from a fixed seed by tools/build_corpus.py
read by id in src/brief.py and src/app.py; never modified after generation
planning notes
data/notes.jsonl -- 40 rows, committed, generated with the cycles
read whole per cycle by src/pack.py; never cached or filtered before the model sees it
gold gap list
data/gold.jsonl -- 40 rows, computed by tools/build_corpus.py and src/segment.py::align at generation time, from a scenario decided before any materiality check runs
never -- evals/scoring.py is pure code, no model, no key
the key
.env -- never committed
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. src/app.py strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 40 cycles, 200 line items across three plan views, the notes log and the gold gap list -- byte-identically from a fixed seed (SEED = 20260818) every time it is run. There is no incremental refresh and no partial rebuild; the materiality bar (src/segment.py::MATERIALITY_PCT, 12%) is a stated constant read fresh from the plan views on every request, not baked into the corpus at generation time.
0.05s wall time to regenerate all 40 cycles, 200 line items and their notes, measured 2026-08-18 -- see Data.index. (lenses.Data.index, tools/build_corpus.py)
a real deployment's plan views change continuously, not on a fixed seed -- how often a real reconciliation cycle needs re-alignment (and whether stale plan views could sit in the notes log alongside a fresher gap) was not measured
point tools/build_corpus.py at your own three plan views and your own line items, and every published completeness/cause figure is void -- they are this corpus's own scenario shapes (see Data.breaks_on), not a property of the model
model
one call per CYCLE carrying every material gap plus the notes log, behind src/adapters/__init__.py, reasoning disabled -- the configuration both scored runs ship
40 of 40 fast-tier and 40 of 40 reasoning-tier cycles returned a reply with at least one gap answered; only 1 of 80 total calls failed to parse at all (GB-0017 on the fast tier, an unterminated JSON string) (lenses.Eval.scores, r001-gap-brief and r002-gap-brief-pro)
the 2,200-token ceiling was never exhausted on either run -- headroom, not a target, per src/prompt.py's own MAX_TOKENS
verdicts are per-model and no configuration ran twice -- the two tiers do not tie: the fast tier beats the reasoning tier sharply on completeness, and the scorer re-runs free on yours
labels
data/gold.jsonl, 40 rows covering 200 line items -- each item's scenario, true cause and (for four of seven scenarios) its exact two explanatory notes lines are decided by tools/build_corpus.py before src/segment.py's materiality check ever runs on the same cycle
146 of 200 line items are gold-material; of those, 99 have a traceable gold cause and 38 have gold cause 'unknown' by design (a missing view or a genuinely unexplained gap) (lenses.Eval.dataset, gap-brief-2026-08-18-40cycles)
your own line items: hand-decide which gaps are real and why, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-adjudicated gold has an error rate this kit has never measured
cause-tag agreement and citation fidelity over this set reflect THIS corpus's five templated cause-phrasings -- a real deployment's planning notes (see Data.bring_your_own_boundary) are untested
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a genuinely-clean line item never appears in a brief, and a genuinely-gapped one always does
the materiality bar (src/segment.py::MATERIALITY_PCT = 12% spread, or a missing view by itself) sat with headroom on both sides of this corpus's planted noise (+-3%) and planted defect shift (18-42%) -- 0 of 200 line items landed ambiguously close to the bar at generation time
if a real deployment's own ordinary noise runs closer to 12% than this corpus's +-3%, re-check MATERIALITY_PCT before trusting which gaps get itemized -- this threshold was never measured against a real distribution (tools/build_corpus.py, src/segment.py, results/eval-r001-gap-brief.json)
a brief's item set for a cycle scores zero true positives despite covering the right COUNT of gaps
the model echoed item_id with the label appended -- 'IT-2 (Base Layers)' instead of 'IT-2' -- which fails the exact-match join for every gap in that cycle at once
check the raw item_id strings in the reply before assuming the model missed the gaps entirely -- 2 of 40 fast-tier cycles and 15 of 40 reasoning-tier cycles on r001/r002 are this exact formatting slip, not a detection failure (results/eval-r001-gap-brief.json and results/eval-r002-gap-brief-pro.json, per_cycle rows with true_positive=0)
a cycle's answer has an empty gaps list and no narrative
the whole reply failed to parse as JSON -- in the one case measured, an unterminated string in the narrative field, with finish_reason 'stop', not a token-budget cutoff
check raw_text (kept only for cycles where gaps_answered < gaps_material) rather than assuming the model saw nothing to report (results/eval-r001-gap-brief.json, cycle GB-0017)
No machine symptom — this failure leaves no trace in any output.
src/pack.py's notes log is sent whole and unfiltered by item, so a model that cites a real line from the wrong item's context has no code check blocking it beyond the item-label-substring test in evals/scoring.py::citation_is_relevant -- a citation phrased without the item's own label (this corpus's templates always include it) would pass the substring-and-relevance check while still being a coincidence, and neither this run nor its scorer can currently tell a deliberate citation from a lucky one.
Concurrency and GPU sizing -- one serial call per cycle, nothing measured past 40. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- Cost prices every call at the cache-miss rate for exactly this reason. Whether a real deployment's own notes-log volume or cause distribution would reproduce this corpus's completeness/cause-tag figures -- see Data.bring_your_own_boundary. Whether either scored run repeats -- every configuration ran once.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The brief's per-item cause tag and citations against that cycle's gold gap list, and its narrative against the packed gap numbers
Catch where a retailer's demand, supply and finance plans disagree
PresenterOpens the private repo. Visible to admins only.
In one lineThe brief's per-item cause tag and citations against that cycle's gold gap list, and its narrative against the packed gap numbers
Did the brief itemize the right set of material gaps, tag each with the right cause (including 'unknown' when gold says unknown), cite two real, item-relevant notes lines whenever it claimed a specific cause, and state only numbers that are actually in the packed gap list?
$0.00per 1,000 planning cycles
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py calls, and the same predicates src/app.py runs on every live draft.
The inputOne real row, seen by every grader
Cycle
GB-0001
Line item
IT-4 -- Rain Shells
Material gap
$44,943.25 spread (32.9%) across demand/supply/financial
Gold cause
assumption_mismatch
The model's cause
assumption_mismatch
Citation 1
Flagging Rain Shells -- confirmed the assumption gap on demand plan is a real difference in what each side assumed, not a data issue.
Citation 2
Rain Shells: demand side is modeling this at a different price/promo assumption than supply side used -- the two aren't reconciled on inputs yet.
Scored as
correct_itemized, correct_cause, citations_real
Grader
Verdict
Why
The brief's per-item cause tag and citations against that cycle's gold gap list, and its narrative against the packed gap numbers
correct_itemized, correct_cause, citations_real
GB-0001, IT-4 (Rain Shells): a 32.9% spread ($44,943.25) across the three plan views. Gold cause: assumption_mismatch. The model independently tagged the same cause, citing two notes lines verbatim -- 'Flagging Rain Shells -- confirmed the assumption gap on demand plan is a real difference in what each side assumed, not a data issue' and 'Rain Shells: demand side is modeling this at a different price/promo assumption than supply side used -- the two aren't reconciled on inputs yet' -- both real, item-relevant lines from that cycle's own notes log.
The formulaWhat it computes
completeness = |brief items ∩ gold material items| over gold (recall) and over brief items (precision); cause-tag agreement = correct causes / matched items, split unknown vs traceable; fabricated_cause = matched items with a non-unknown cause and a citation that is not both real and item-relevant; narrative faithfulness = cycles with zero unmatched quantified claims / cycles with a narrative.
The analysisWhat it actually did
Model
Result
the fast tier
93.8% gap completeness recall · 3 more measured on this row
the reasoning tier
64.4% gap completeness recall · 3 more measured on this row
In operationWhat to monitor
Reference standard: The corpus's own gold cause and citations, decided by tools/build_corpus.py at generation time, before src/segment.py's materiality check runs on the same cycle. This grader IS the reference, so its own TPR/TNR are not separately measured -- that would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and largely cannot be -- it IS the reference. What can go wrong is the corpus's own planted causes and citations, which come from the generator's fixed templates.
Watch these
dropped_gap
fabricated_gap
wrong_cause
fabricated_cause
Alarm on
Any fabricated_cause on any run -- the guardrail this kit's own atlas facet sheet names first. Both scored tiers had zero.
How tight can the band be? There is no tuned threshold anywhere in this grader: materiality is a stated constant in src/segment.py, and the citation check is a literal substring-and-relevance test, not a fitted score.
Cadence: Re-run on any change to src/prompt.py, src/segment.py, src/rubric.py or the corpus generator.
The decisionWhen to reach for it
Use it
The gold cause and gold citations are known -- true of every kit corpus, never true of a real deployment's own planning cycle.
Do not use it
The truth is not known -- the normal state of a real reconciliation meeting, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning