Tell a vendor exactly where their invoice is stuck
Vendors keep asking when an invoice will be paid, and the answer is spread across match, approval, payment run and payment. This app reads all four in order, finds the step really holding the invoice, and drafts the reply.
PresenterOpens the private repo. Visible to admins only.
For accounts payableCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
Accounts payable staff at a retailer or any business with many suppliers, answering vendor emails about payment.
✕Today's manual process
1Pull up the invoice and read its match, approval, payment run and payment status.
2Work out which step is holding it, even when a later field looks finished.
3Write the vendor a reply and decide whether AP should check it before it goes out.
4One wrong "paid" tells a vendor money is coming while the invoice is still stuck.
Every invoice checked manually, step by step
✓With the app
1The whole record is read in order: match, then approval, then payment run, then payment.
2The step really holding it is named, even when a later field looks finished.
3The reply is drafted from that step, and flagged for AP review when a problem is still open.
4A reply says "paid" only when the record shows the payment went out. Nothing here releases or moves money.
Replies drafted, open problems flagged
See it work
One real case: what the app reads, step by step
Larkspur Catering Group asks about payment timing on a $49,904.91 invoice that matched but is escalated at approval, over the manager's limit.
Tell a vendor exactly where their invoice is stuckReference appBuilt to be shaped to your process
6
1The invoice the record the app reads: vendor, purchase order, amount and the date it came in.
2Match done: the invoice matched, so the first step is clear.
3Approval, the first step not finished the amount is over the manager's limit, so it was escalated.
4Not yet, downstream run inclusion and remittance both wait until approval is complete.
5The vendor's question asking for any update on payment timing.
6The outcome held at approval, AP review required, and the reply drafted from that step.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Tell a vendor exactly where their invoice is stuck
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Reconciling a vendor's payment-status inquiry means reading the invoice's full match/approval/run-inclusion/remittance trail, working out which of the four stages is the real bottleneck -- never trusting whatever a downstream field happens to show -- and drafting a reply grounded in that stage, flagged for AP review whenever an exception is still open, for every inquiry, every invoice. AP staff reading a vendor's payment-status inquiry against its full match/approval/run-inclusion/remittance trail by hand, working out which of the four stages actually governs before drafting a reply -- for every inquiry, every invoice.
Audience
AP staff who answer vendor payment-status inquiries, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual invoices
The corpus is 35 invoices, 0.02 MB (jsonl 1). A real vendor payment run names a real vendor, a real bank reference and a real approval chain -- this repo has never held one and is not able to. There is no public corpus of paired (invoice 4-stage trace, vendor inquiry, adjudicated current stage) records, for the same reason no public corpus of production alerts or supplier agreements exists: the interesting cases are exactly the ones nobody can publish. Nothing in the corpus refers to a real company -- vendor names, invoice numbers, PO numbers and amounts are all invented.
The corpus
The 35 invoicesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your invoices. That is the whole change — there is no database to migrate.
One invoice, as the model receives itinvoices.jsonl · 1 of 35
{"invoice_id": "INV-2024-0600", "vendor_name": "Larkspur Catering Group", "po_number": "PO-31257", "amount": 49904.91, "submitted_date": "2024-07-02", "inquiry": "Following up on INV-2024-0600 -- any update on payment timing?", "match": {"matched": true, "reason": null}, "approval": {"status": "exception", "reason": "amount exceeds manager approval threshold, escalated"}, "run_inclusion": {"included": false, "run_id": null, "scheduled_date": null, "reason": "not yet eligible for run inclusion; approval not complete"}, "remittance": {"remitted": false, "remittance_date": null, "method": null, "reference": null, "reason": "not yet remitted; approval not complete"}}
{"invoice_id": "INV-2024-0853", "vendor_name": "Northgate Logistics", "po_number": "PO-16783", "amount": 30450.06, "submitted_date": "2024-06-07", "inquiry": "Following up on INV-2024-0853 -- any update on payment timing?", "match": {"matched": true, "reason": null}, "approval": {"status": "exception", "reason": "invoice amount exceeds the original purchase order by more than tolerance, escalated for review"}, "run_inclusion": {"included": false, "run_id": null, "scheduled_date": null, "reason": "not yet eligible for run inclusion; approval not complete"}, "remittance": {"remitted": false, "remittance_date": null, "method": null, "reference": null, "reason": "not yet remitted; approval not complete"}}
Abridged — the file continues.
The outcomeWhat a good result looks like
A traced current stage, an AP-review flag derived from it, the one real date (if any) the record supports, and a drafted reply grounded in that stage -- informational only, never releasing a payment or committing to a date beyond the invoice's own scheduled run.
And when it cannot
Telling a vendor an invoice is paid -- either stating current_stage=remitted or a reply that otherwise claims payment happened or is imminent -- when the invoice's true governing stage is actually an open match or approval exception. Measured at 0 of 4 planted remittance-trap invoices this run -- see not_good_enough for why zero on one run is not the same claim as zero on a larger or adversarial sample.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Tracing a vendor's payment-status inquiry against its full match/approval/run-inclusion/remittance trail and drafting a grounded reply — the fast tier, over the free rule-order floor 100% stage accuracy and 0% false_paid this run, against the floor's 82.9% and 100% false_paid rate on the exact invoices built to be misread downstream-first.
Deciding whether a reply can state a payment has actually gone out — the fast tier's current_stage verdict, read alongside whether the reply text itself makes a paid claim 0 of 4 gold remittance-trap invoices this run were told to a vendor as paid (false_paid=0) -- the expensive-direction error this kit's guardrail exists to catch.
What this kit is not — an informational status reply for AP to review before it reaches a vendor -- never an approval or a payment release src/payrun.py and src/app.py have no function anywhere that releases a payment or changes a remittance detail; every reply with an open exception is flagged requires_ap_review before it goes out.
At a glanceHow the whole thing runs
100%stage accuracy pct
2,629 msp50, end to end
$1.38per 1,000 invoices · Google Gemini 3 Flash
Run once, for real, on 2026-08-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Tell a vendor exactly where their invoice is stuck14 steps · 4 questions · run once, for real · 2026-08-20
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own vendors, amounts and stage reasons, or write your own data/invoices.jsonl in the same shape -- src/payrun.py and src/prompt.py read an invoice by its own fields (match, approval, run_inclusion, remittance, inquiry) and do not care where the values came from. The measured 100% stage accuracy and 0% false_paid rate are THIS corpus's inquiry phrasing (a handful of hand-written templates) and THIS corpus's one-trap-family-per-invoice design (TRAP_N=6 of 14 exception-stage invoices, ~43%, split 4 remittance / 2 run_inclusion).Corpus lens →
When is this the wrong choice?
Avoid: The rule-order floor for anything beyond the 29 non-trap invoices, where checking remittance first happens to agree with the correct precedence -- it is not a competitor, it is the honest floor a model has to clear. That is the case against the best-fitting scenario (“Tracing a vendor's payment-status inquiry against its full match/approval/run-inclusion/remittance trail and drafting a grounded reply”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
This kit assumes match, approval and run-inclusion status all live in systems it can query directly -- a real deployment where an approval step is tracked outside those systems (an email sign-off, a verbal override) has no trace to follow, and this kit has no way to notice the trail it is reading is incomplete rather than merely unfavourable. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether this run's 100% stage accuracy and 0 false_paid are stable properties of this model at these settings, or a one-off -- one recorded run is not a history. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-20 — r001-fin-payrun. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- every invoice's full 4-stage record and its vendor inquiry -- since data/invoices.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Clicking Trace with no API_KEY returns a calm 200 explaining nothing was called (src/app.py's /api/check). It cannot reproduce a live trace, a score, or a dollar figure without a key -- those are what results/eval-r001-fin-payrun.json already committed.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
2,629 msp50, end to end
4,190 msp95
2 minclone to first result
What the clock covers. END-TO-END per vendor inquiry: one HTTP request carrying the invoice's full 4-stage record and the vendor's inquiry text, parsed to a traced current stage, an AP-review flag, a stated date and a drafted reply. No retrieval step -- loading an invoice from disk is pure code and costs no network time. This run left provider-side reasoning ('thinking') at its documented default (on) rather than disabling it, so p50/p95 include a hidden reasoning pass on every call, not only the JSON-formatting cost -- see not_good_enough.
Current processWhat it replaces
AP staff reading a vendor's payment-status inquiry against its full match/approval/run-inclusion/remittance trail by hand, working out which of the four stages actually governs before drafting a reply -- for every inquiry, every invoice.
Where it is not good enough
Three real findings. First, a live discrepancy between what this report measures and what a forker's own click actually does: the registered run (r001-fin-payrun) left provider-side reasoning ('thinking') at its documented default -- ON -- while the shipped app's own endpoint (src/app.py's /api/check) hardcodes thinking=THINKING_OFF on every live call. Nobody reconciled the two before this run was registered as the measured fact. Reasoning consumed 72.7% of this run's total output-token budget (6,960 of 9,575 output tokens; 198.9 tokens/call average, up to 925 on one invoice) and is a real driver of both this kit's latency (p50 2,629ms, p95 4,190ms) and, since a provider bills reasoning tokens as completion tokens, its dollar cost -- the published cost and latency numbers on this report are therefore NOT what a reader gets from the shipped UI, that configuration has not been separately measured. Second, this run's own 100% stage accuracy and 0% false_paid rate are measured on one 35-invoice run, not a distribution -- no repeat exists (see Eval.could_not_verify), and false_paid's own denominator is exactly 4 planted remittance-trap invoices, all correctly traced this run, which is what a corpus of this scale honestly supports (see data/SOURCES.md) rather than a stable rate over a larger sample. Third, a structural limitation stated in the corpus's own documentation, not found by this run: this kit assumes match, approval and run-inclusion status all live in systems it can query directly -- a manually tracked approval step (an email sign-off that never made it into the approval system) has no trace to follow, and this kit has no way to notice that the trail it is reading is incomplete rather than merely unfavourable; every invoice in this corpus has a complete trail by construction.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free rule-order baseline (evals/baseline.py, a fixed-priority rule that checks remittance before match/approval/run_inclusion) scores 82.9 pct stage accuracy overall but misclassifies all 6 planted trap invoices — a 100 pct false_paid rate on the 4 remittance-trap invoices — exactly the downstream-field trap this kit's model resolved correctly on all 6 planted instances. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
whether reasoning is enabled
src/payrun.py and src/app.py
check() takes a thinking kwarg; src/app.py's live /api/check hardcodes it OFF while the registered eval run (r001-fin-payrun) left it at provider default (on) -- a real, live discrepancy between what this report measures and what a forker's own click experiences today. See Business.not_good_enough.
the stage taxonomy and precedence order
src/prompt.py
A stage this kit's fixed five-value STAGES tuple does not carry, or a different read order, needs a rewrite of SYSTEM plus REVIEW_STAGES -- a prompt tweak alone will not teach a precedence rule the corpus never plants.
the corpus
tools/build_corpus.py
Point it at your own vendors, amounts and stage reasons. src/payrun.py and src/prompt.py read an invoice by its own fields (match, approval, run_inclusion, remittance, inquiry) and do not care where the numbers came from.
Set by three module constants, not a single global -- a new corpus states its own trap counts in the same three names.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 35 invoice records and vendor inquiries from a fixed seed (SEED=20260819). Gold (current_stage, requires_ap_review, expected_date) is RECOMPUTED from the actual generated 4-stage record by derive_stage(), never from the target_stage the generator aimed for -- --verify independently re-checks every row against the record actually written to disk.
the prompt
src/prompt.py
The five-stage taxonomy, the read-in-order precedence rule and the answer schema, declared once and read from here by build() and parse(). States the false-'remitted' failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover.
the AI layer
src/payrun.py
Loads one invoice by id, calls the model once with its full 4-stage record and the vendor inquiry, parses the reply into a stage, a review flag, a stated date and a drafted reply. The whole AI layer, deliberately short -- the same split fin-invval's invval.py and fin-close's close.py both make. Never releases a payment or changes a remittance detail -- there is no function here or in src/app.py that does either.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
spend control
src/budget.py
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
the app
src/app.py
Static files and two JSON endpoints (/api/invoice, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-fin-payrun) actually used; see Business.not_good_enough.
the scorer
evals/scoring.py
Stage-level scoring, pure code: exact match on current_stage and requires_ap_review, false_paid (a keyword scan of the reply plus a remitted-stage check) reported as its own named figure rather than folded into a generic 'wrong' bucket, and date accuracy checked separately.
free baseline
evals/baseline.py
A fixed-priority rule, written the way a person free-texting a quick classifier would write one, that deliberately checks remittance BEFORE match/approval/run_inclusion -- so it misclassifies every planted trap invoice by construction. The honest, narrow floor a model has to clear.
Where it breaks at scale
Not on invoice count -- each invoice is independent and one call per invoice is linear. It breaks two other ways. First, the STRUCTURAL ASSUMPTION: this kit assumes match, approval and run-inclusion status all live in systems it can query directly (see Data.breaks_on) -- a manually tracked approval step (an email sign-off never entered into the approval system) has no trace to follow, and this kit has no way to notice the trail is incomplete rather than merely unfavourable; stated plainly in data/SOURCES.md, not discovered by this run. Second, REASONING LEFT ON NEAR THE CEILING: MAX_TOKENS is fixed at 3000 (src/payrun.py), sized generously up front rather than computed -- this run's worst observed call used 1,017 of the 3,000-token budget (925 of it reasoning), leaving real headroom, but reasoning consumed 72.7% of the average output-token budget and was left at provider default (on) rather than disabled -- a call with meaningfully more reasoning than this run's worst case was never measured against the ceiling.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Invoice INV-2024-0600 -- Larkspur Catering Group, PO-31257, $49,904.91: the full 4-stage record shown (match clean, approval exception, run inclusion and remittance both not yet reached) against the vendor's inquiry, before any call is made.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page after pressing Trace with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. No traced stage or reply is shown -- this is the honest failure state, not a staged one.failureOpen full size →
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
35invoices
0.02 MiBjsonl 1
p50 410chars per characters per invoice's assembled record block
$0.00setup · 0.0s
How it is cutWhat one characters per invoice's assembled record block is
One invoice, its full 4-stage record and one vendor inquiry included whole -- no train/test split, every invoice traced once, in one call.
SetupWhat the setup figure measured
No index is built. src/app.py finds an invoice by scanning C.invoices() for a matching invoice_id -- 35 rows, a linear scan, not a search structure -- and src/payrun.py sends that invoice's full record whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own invoices
Point tools/build_corpus.py at your own vendors, amounts and stage reasons, or write your own data/invoices.jsonl in the same shape -- src/payrun.py and src/prompt.py read an invoice by its own fields (match, approval, run_inclusion, remittance, inquiry) and do not care where the values came from. The five-stage taxonomy in src/prompt.py assumes exactly these four gating systems (match, approval, run_inclusion, remittance); a different set of gating stages is a prompt.py change, not a data change.
⚠︎ And what stops being true when you do: The measured 100% stage accuracy and 0% false_paid rate are THIS corpus's inquiry phrasing (a handful of hand-written templates) and THIS corpus's one-trap-family-per-invoice design (TRAP_N=6 of 14 exception-stage invoices, ~43%, split 4 remittance / 2 run_inclusion). A real vendor's own inquiry phrasing, and a real payment system's own trap shapes -- including ones this corpus does not model, like an incomplete rather than unfavourable trail (see breaks_on) -- are both unmeasured by this run.
What breaks it
This kit assumes match, approval and run-inclusion status all live in systems it can query directly -- a real deployment where an approval step is tracked outside those systems (an email sign-off, a verbal override) has no trace to follow, and this kit has no way to notice the trail it is reading is incomplete rather than merely unfavourable.
One planted trap family per invoice, at most: a downstream field (remittance, or less often run_inclusion) that looks complete while an earlier stage governs. A real payment run can carry a stale-looking field at more than one stage on the same invoice at once; this corpus plants at most one to keep the failure mode legible.
The trap denominator is small. 4 remittance-trap and 2 run_inclusion-trap invoices is what a corpus of this scale supports honestly -- a wider real deployment would want a larger sample before trusting the false_paid rate as a stable number.
One invoice, one static snapshot -- not a running history. A real deployment would see the same invoice recur across multiple inquiries as its stages change; this corpus judges every invoice independently, once.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,634
989
record
481
131
Total
1,120
This is the cost lesson as arithmetic: of the 1,120 tokens assembled, 989 are instructions — 88% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged verbatim by importing src.prompt and src.payrun directly and calling prompt.build(inv) for INV-2024-0600 (not retyped) -- the literal system and user content src/payrun.py::check() sends. Per-part token counts are a proportional estimate over character share (system 3634 / record 481 chars, 4115 total) applied to this call's own total of 1120 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split. The 'record' part covers only the invoice's 4-stage block and inquiry line src/prompt.py::_stage_block() builds -- the surrounding 'Trace the invoice below...' instruction text is part of the user message actually sent (see prompt_verbatim) but is not a separately named segment in the kit's own code.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You trace one vendor invoice through match, approval, payment-run inclusion and remittance, and draft a reply to the vendor's payment-status inquiry grounded in its actual traced state. You are given the invoice's full 4-stage record and the vendor's inquiry text.
THE FOUR STAGES, READ IN THIS ORDER, STOPPING AT THE FIRST ONE THAT IS NOT CLEANLY COMPLETE
1. match matched true/false, with a reason if false
2. approval status approved/pending/exception, with a reason if exception
3. run_inclusion included true/false; if true a run_id + scheduled_date, if false a reason
4. remittance remitted true/false; if true a remittance_date + method + reference, if false a reason
The TRUE current stage is whichever of these four is the actual bottleneck, read IN ORDER -- match first, then approval, then run_inclusion, then remittance. Stop at the first stage that is not cleanly complete. This is the trap: a later field can look complete -- a run_inclusion with a real run_id and scheduled_date, a remittance with a real date and reference -- while an EARLIER stage is the one that actually governs. A downstream field that looks done never overrides an earlier stage that is not. Read the whole record before answering; do not stop at the first field that looks reassuring.
THE FIVE VALUES FOR current_stage, EXACTLY ONE APPLIES
match_exception match.matched is false -- the invoice has not cleared three-way match and nothing downstream can be trusted regardless of what it shows.
approval_exception match is clean, and approval.status is "exception" -- the invoice is held on an approval problem, whatever run_inclusion or remittance happen to contain.
awaiting_run_inclusion match and approval are both clean, and run_inclusion.included is false -- approved, but not yet placed on a scheduled payment run.
in_scheduled_run match, approval and run_inclusion are all clean, and remittance.remitted is false -- on a scheduled run with a real scheduled_date, not yet paid.
remitted all four stages are clean and remittance.remitted is true -- the payment has actually gone out.
requires_ap_review is true if and only if current_stage is match_exception or approval_exception -- an open exception at either stage means the drafted reply needs AP review before it is sent to the vendor.
THE REPLY MUST BE GROUNDED IN current_stage, NEVER IN A DOWNSTREAM FIELD ALONE
- Never claim the invoice is paid, or that payment is imminent, unless current_stage is remitted.
- If current_stage is in_scheduled_run, state ONLY the real scheduled_date from run_inclusion -- never an earlier or invented date, never "accelerated" or "off-cycle" language.
- If current_stage is remitted, state the real remittance_date and reference.
- If current_stage is match_exception or approval_exception, the reply must not state a payment timeline at all -- there is nothing to schedule yet -- and must say the item needs AP review.
stated_date is the one date, if any, the reply states: the real scheduled_date for in_scheduled_run, the real remittance_date for remitted, or null for the other three stages -- there is no date to state for an exception or for a stage still awaiting run inclusion.
You never release a payment, never change a remittance detail, and never commit to a date beyond the invoice's own scheduled run. The reply is informational only -- it reports the invoice's actual traced state to the vendor, it never promises to accelerate or move a payment off-cycle, and it is never sent without AP review when requires_ap_review is true.
Trace the invoice below and draft a reply to the vendor's inquiry.
Return a JSON object with exactly four keys: "current_stage" (one of: match_exception, approval_exception, awaiting_run_inclusion, in_scheduled_run, remitted), "requires_ap_review" (true or false), "stated_date" (the one date the reply states, as it appears in the record, or null), "reply" (the drafted reply text, one or two sentences).
INVOICE INV-2024-0600 (vendor Larkspur Catering Group, PO PO-31257, amount $49904.91, submitted 2024-07-02)
1. match: matched=True
2. approval: status=exception reason: amount exceeds manager approval threshold, escalated
3. run_inclusion: included=False reason: not yet eligible for run inclusion; approval not complete
4. remittance: remitted=False reason: not yet remitted; approval not complete
VENDOR INQUIRY: Following up on INV-2024-0600 -- any update on payment timing?
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"current_stage": "approval_exception",
"requires_ap_review": true,
"stated_date": null,
"reply": "Your invoice INV-2024-0600 is currently held for AP review because it exceeds the manager approval threshold and has been escalated. We cannot provide a payment timeline until the approval exception is resolved."
}
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Tell a vendor exactly where their invoice is stuck — 35 invoices. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: exact match on current_stage (the core 5-way classification) and on requires_ap_review (its own field even though it is derived deterministically from current_stage in gold, because the model can get the stage right and the boolean wrong, or the reverse, if it does not apply the derivation itself); false_paid (a keyword scan of the drafted reply plus a current_stage=='remitted' check) counted separately as the expensive-direction error, over exactly the corpus's planted remittance-trap denominator; and date accuracy (does stated_date match the one real date the record supports, or correctly state none). No judge model. The same function scores both evals/baseline.py's free floor and evals/run.py's real run.
35invoices
35source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED35 / 35stage accuracy pct — invoices with a parsed stage answerDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED35 / 35stage answered pct — invoices askedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED35 / 35review accuracy pct — invoices with a parsed AP-review answerDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 4false paid rate pct — planted remittance-trap invoicesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED35 / 35date accuracy pct — invoices scored for stated dateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/scoring.py) is pure code, exact match against a mechanically-derived gold stage -- there is no judgement to validate, only string/boolean comparison and a keyword regex. What WAS validated: tools/build_corpus.py's gold current_stage, requires_ap_review and expected_date are all RECOMPUTED by derive_stage() from the same generated 4-stage record the model reads, never carried over from the target_stage the corpus builder aimed for -- and --verify independently re-derives every row from data/invoices.jsonl on disk and asserts zero drift, confirmed this session (35 invoices checked, 0 drift; 6 trap invoices, 4 remittance / 2 run_inclusion, all resolving as documented).
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One invoice
1,000 invoices
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.001381
$1.38
41%
Same work, 1× the bill
The same invoices, the same tokens — only the rate card changed. And on that card about 41% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether reasoning ('thinking') is left on or explicitly disabled for this model -- the app's own /api/check always disables it; the registered run left it at provider default. The two configurations have not been priced against each other on this corpus.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page. Reasoning was also left at provider default (on) for this run -- see Business.not_good_enough -- so even the projected figure prices tokens the shipped app's own reasoning-off configuration would not spend.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing -- exact match on stage/review/date and a keyword scan for a paid claim in the reply are both pure code. A measured $0.00, not an unpriced one.
The gradersOne way to grade, and why it is the only one
The floor checks remittance FIRST, so any invoice whose remittance field shows remitted:true is called remitted regardless of what match or approval actually say -- by construction it misclassifies all 4 gold remittance-trap invoices as remitted (a 100% false_paid rate, 4 of 4) and the 2 run_inclusion-trap invoices as in_scheduled_run (checked before approval), for 6 total misses -- exactly the corpus's planted trap set (results/eval-b000-ruleorder.json's own misses list). Its overall stage accuracy still reaches 82.9% (29 of 35) because the other 29 invoices have nothing downstream that looks complete while an earlier stage governs, so both orders agree on them. The fast tier's 0% false_paid and 100% stage accuracy on the identical 35 invoices is the real, measured gap a model closes that this floor cannot.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Stage precedence exact match, plus AP-review derivation, false-paid claim scan and date exact match Does the model's current_stage match gold's precedence-derived stage? Does requires_ap_review match the stage-derived boolean? Does the reply, or the stage itself, falsely claim the invoice is paid when a gold remittance-trap invoice's true stage is still an open exception? Does stated_date match the one real date the record supports, or correctly state none?
$0.00
no
yes
the fast tier 100.0% stage accuracy · 1 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, and cleanly on the one axis this corpus was built to test: the fast tier and the free rule-order floor tie on the 29 non-trap invoices (both orders agree when nothing downstream looks complete while an earlier stage governs), then diverge sharply on the 6 planted trap invoices -- 100% (fast tier) vs 0% (floor, i.e. the floor gets every one wrong) -- because the floor's remittance-first check answers exactly what the trap is built to make it answer, and the fast tier does not. A grader that could not tell the two apart would not produce a gap that lines up this precisely with the corpus's own documented trap counts.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Tracing a vendor's payment-status inquiry against its full match/approval/run-inclusion/remittance trail and drafting a grounded reply
the fast tier, over the free rule-order floor
100% stage accuracy and 0% false_paid this run, against the floor's 82.9% and 100% false_paid rate on the exact invoices built to be misread downstream-first.
the rule-order floor for anything beyond the 29 non-trap invoices, where checking remittance first happens to agree with the correct precedence -- it is not a competitor, it is the honest floor a model has to clear.
Deciding whether a reply can state a payment has actually gone out
the fast tier's current_stage verdict, read alongside whether the reply text itself makes a paid claim
0 of 4 gold remittance-trap invoices this run were told to a vendor as paid (false_paid=0) -- the expensive-direction error this kit's guardrail exists to catch.
trusting a remitted verdict without checking whether reasoning was left on for the call that produced it -- see Business.not_good_enough; this run's own configuration has not been re-measured with reasoning explicitly disabled the way the live app runs it.
What this kit is not
an informational status reply for AP to review before it reaches a vendor -- never an approval or a payment release
src/payrun.py and src/app.py have no function anywhere that releases a payment or changes a remittance detail; every reply with an open exception is flagged requires_ap_review before it goes out.
sending this kit's drafted reply straight to a vendor with no AP review step for any invoice flagged requires_ap_review.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
match_exception
match.matched is false -- the invoice has not cleared three-way match
6
INV-2024-0842: match.matched=false ("duplicate invoice number flagged"). Gold and model both: match_exception. Model reply: "Our records show invoice INV-2024-0842 is currently under AP review because it was flagged as a possible duplicate. We will provide an…
approval_exception
match clean, approval.status is exception
8
INV-2024-0600: match clean, approval.status=exception ("amount exceeds manager approval threshold, escalated"). Gold and model both: approval_exception.
awaiting_run_inclusion
match and approval clean, run_inclusion.included is false
6
INV-2024-0548: approved, run_inclusion.included=false ("on hold pending a repeat three-way match re-verification"). Gold and model both: awaiting_run_inclusion.
in_scheduled_run
match, approval and run_inclusion clean, remittance.remitted is false
8
INV-2024-0808: on RUN-2024-014, scheduled 2024-07-31, remittance.remitted=false. Gold and model both: in_scheduled_run; reply states only the real scheduled_date.
remitted
all four stages clean and remittance.remitted is true
7
INV-2024-1145: remittance.remitted=true, remittance_date=2024-04-11, ACH, REM-977401. Gold and model both: remitted.
remittance_trap_resolved
Invoices whose remittance field shows a complete-looking payment while an earlier stage (match or approval) is still the true bottleneck, resolved correctly
4
INV-2024-1020: approval.status=exception (amount exceeds PO tolerance) but remittance shows remitted=true, remittance_date=2024-08-06, wire, REM-786773 -- a real payment record for an invoice later pulled back for approval. Gold approval_exception. The…
run_inclusion_trap_resolved
Invoices whose run_inclusion field shows a real scheduled run while approval is still the true bottleneck, resolved correctly
2
INV-2024-1940: approval.status=exception ("budget code mismatch, held for correction") but run_inclusion shows included=true, RUN-2024-035, scheduled 2024-06-28 -- a real scheduled run the invoice was later pulled from pending resolution. Gold…
What we could NOT verify
Whether this run's 100% stage accuracy and 0 false_paid are stable properties of this model at these settings, or a one-off -- one recorded run is not a history.
Whether the same figures hold with reasoning explicitly disabled. The registered run (r001-fin-payrun) left provider-side reasoning at its documented default (on); the shipped app's own /api/check hardcodes it off. No run exists at the app's actual setting -- see Business.not_good_enough.
Whether resistance holds against an adversarial or malicious vendor inquiry -- no red-team run exists for this kit (see redteam_page).
Whether a real payment-run trail with a genuinely incomplete step (an approval sign-off that never made it into the tracked system) would be detected as incomplete rather than misread as an exception or a clean stage -- see Data.breaks_on.
Whether a real vendor's own inquiry phrasing (rather than this corpus's handful of templates) would reproduce these figures -- see Data.bring_your_own_boundary.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
1,120.0
273.57
2,629 ms
$0.001381
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The rule-order baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 39,200 input / 9,575 output tokens priced at Google Gemini 3 Flash's published rate -- the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The invoice's 4-stage record (match, approval, run_inclusion, remittance) plus the vendor inquiry text -- fixed size per call by construction (input tokens ranged 1,103-1,142 across all 35 calls, essentially flat), the floor every call pays regardless of which stage governs.
Reasoning ('thinking') was left at the provider's default (on) rather than explicitly disabled for this run -- it consumed 72.7% of the average output-token budget (198.9 of 273.6 tokens/call) and is already priced into cost_per_query_usd, since a provider bills reasoning tokens as completion tokens. src/app.py's live UI hardcodes thinking off, a different, unmeasured configuration -- see Business.not_good_enough.
The fixed system prompt (3,634 characters -- the stage taxonomy, precedence rule and answer schema) is sent in full on every call, the largest single fixed cost per invoice.
Your volumeWhat it costs at your volume
Linear in invoices: each call is independent and self-contained, with no shared context or retrieval step to amortise. This run's 35 invoices cost about $0.0483 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $0.483 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volumevendor payment-status inquiries traced per day/week -- this run traced 35 in one pass
What it replacesAP staff reading a vendor inquiry against the invoice's full match/approval/run-inclusion/remittance trail by hand and drafting a reply grounded in the real governing stage
Time saved per itemnot measured here -- depends on how long a manual trace-and-reply pass takes at the reader's own company
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run -- unlike gap-brief's fast/reasoning comparison -- so this page prices one model, not a trade-off; see Eval.could_not_verify for what a second run would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
39,200input tokens · this run
9,575output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 35 vendor inquiries traced, 35 invoices stage-scored by pure code. This run (r001-fin-payrun) answered 35 of 35 stages and 35 of 35 AP-review flags -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.019
$0.019
$0.55
2026-09-12
gemini-3-flash
Google
$0.048
$0.048
$1.38
2026-09-18
gemini-3-8-flash
Google
$0.065
$0.065
$1.87
2026-09-18
claude-haiku-4-5
Anthropic
$0.087
$0.087
$2.49
2026-09-12
llama-5
Meta
$0.090
$0.090
$2.56
2026-09-18
grok-4-5
xAI
$0.136
$0.136
$3.88
2026-09-18
grok-4-6
xAI
$0.136
$0.136
$3.88
2026-09-18
claude-sonnet-5
Anthropic
$0.174
$0.174
$4.98
2026-09-12
gemini-3-1-pro
Google
$0.193
$0.193
$5.52
2026-09-18
gpt-5-6-terra
OpenAI
$0.193
$0.193
$5.52
2026-09-12
gpt-5-6-sol
OpenAI
$0.348
$0.348
$9.95
2026-09-12
claude-opus-4-8
Anthropic
$0.435
$0.435
$12.44
2026-09-12
claude-opus-5
Anthropic
$0.435
$0.435
$12.44
2026-09-12
claude-fable-5
Anthropic
$0.871
$0.871
$24.88
2026-09-18
claude-fable-5-1
Anthropic
$0.871
$0.871
$24.88
2026-09-18
gpt-6-astra
OpenAI
$0.871
$0.871
$24.88
2026-09-17
Read this against the numbers above
REASONING WAS NOT DISABLED FOR THIS WORKLOAD. Every row below prices THIS run's own token counts, which include reasoning tokens the live app's own /api/check never generates (it hardcodes thinking off). A forker's real bill on other models depends on whether that model has an equivalent reasoning toggle and whether it is left on.
NO QUALITY IS IMPLIED. Only the model that produced Eval.scores has been scored against this corpus -- every row here is a price, not a recommendation.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 35 invoice records and vendor inquiries from a fixed seed (SEED=20260819). Gold (current_stage, requires_ap_review, expected_date) is RECOMPUTED from the actual generated 4-stage record by derive_stage(), never from the target_stage the generator aimed for -- --verify independently re-checks every row against the record actually written to disk.
You change it to: Point it at your own vendors, amounts and stage reasons. src/payrun.py and src/prompt.py read an invoice by its own fields (match, approval, run_inclusion, remittance, inquiry) and do not care where the numbers came from.
tools/build_corpus.py
# Generate the invoice records and vendor inquiries this kit checks against, from a fixed seed.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
STAGE_COUNTS = {
TRAP_N = 6
TRAP_REMITTANCE_N = 4
TRAP_RUN_INCLUSION_N = 2
VENDOR_NAMES = [
MATCH_REASONS = [
src/prompt.pythe prompt — a swap seam
The five-stage taxonomy, the read-in-order precedence rule and the answer schema, declared once and read from here by build() and parse(). States the false-'remitted' failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover.
You change it to: A stage this kit's fixed five-value STAGES tuple does not carry, or a different read order, needs a rewrite of SYSTEM plus REVIEW_STAGES -- a prompt tweak alone will not teach a precedence rule the corpus never plants.
src/prompt.py
# Assemble the one prompt this kit sends, and parse the one reply it gets back.
STAGES = (
STAGE_MEANINGS = {
REVIEW_STAGES = ("match_exception", "approval_exception")
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _stage_block(inv):
def build(inv, prompt=DEFAULT_PROMPT):
def parse(raw):
src/payrun.pythe AI layer
Loads one invoice by id, calls the model once with its full 4-stage record and the vendor inquiry, parses the reply into a stage, a review flag, a stated date and a drafted reply. The whole AI layer, deliberately short -- the same split fin-invval's invval.py and fin-close's close.py both make. Never releases a payment or changes a remittance detail -- there is no function here or in src/app.py that does either.
src/payrun.py
# Trace one invoice through match, approval, run inclusion and remittance, and draft a reply to
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
INVOICES = os.path.join(HERE, "data", "invoices.jsonl")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
MAX_TOKENS = 3000
def invoices():
def load_gold():
def check(cfg, inv, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pyspend control
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
src/app.pythe app
Static files and two JSON endpoints (/api/invoice, /api/check) on the standard library. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-fin-payrun) actually used; see Business.not_good_enough.
src/app.py
# The minimal local UI. Standard library only — python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8789"))
class H(BaseHTTPRequestHandler):
def main():
evals/scoring.pythe scorer
Stage-level scoring, pure code: exact match on current_stage and requires_ap_review, false_paid (a keyword scan of the reply plus a remitted-stage check) reported as its own named figure rather than folded into a generic 'wrong' bucket, and date accuracy checked separately.
evals/scoring.py
# Score a set of predicted invoice traces against gold. Pure code, shared by evals/baseline.py
STAGES = ("match_exception", "approval_exception", "awaiting_run_inclusion", "in_scheduled_run",
def _reply_claims_paid(reply):
def score(records, gold):
evals/baseline.pyfree baseline
A fixed-priority rule, written the way a person free-texting a quick classifier would write one, that deliberately checks remittance BEFORE match/approval/run_inclusion -- so it misclassifies every planted trap invoice by construction. The honest, narrow floor a model has to clear.
evals/baseline.py
# What a simple, dumb rule catches, over the same corpus. Free. No key, no dependency, no model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def classify(inv):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 35 invoice records and vendor inquiries from a fixed seed (SEED=20260819). Gold (current_stage, requires_ap_review, expected_date) is RECOMPUTED from the actual generated 4-stage record by derive_stage(), never from the target_stage the generator aimed for -- --verify independently re-checks every row against the record actually written to disk. A swap seam.
src/prompt.pyThe five-stage taxonomy, the read-in-order precedence rule and the answer schema, declared once and read from here by build() and parse(). States the false-'remitted' failure mode to the model explicitly, as the named failure this task exists to catch, not left for the model to discover. A swap seam.
src/payrun.pyLoads one invoice by id, calls the model once with its full 4-stage record and the vendor inquiry, parses the reply into a stage, a review flag, a stated date and a drafted reply. The whole AI layer, deliberately short -- the same split fin-invval's invval.py and fin-close's close.py both make. Never releases a payment or changes a remittance detail -- there is no function here or in src/app.py that does either.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from. A swap seam.
src/budget.pyAn append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose -- see the file's own docstring.
evals/scoring.pyStage-level scoring, pure code: exact match on current_stage and requires_ap_review, false_paid (a keyword scan of the reply plus a remitted-stage check) reported as its own named figure rather than folded into a generic 'wrong' bucket, and date accuracy checked separately.
evals/baseline.pyA fixed-priority rule, written the way a person free-texting a quick classifier would write one, that deliberately checks remittance BEFORE match/approval/run_inclusion -- so it misclassifies every planted trap invoice by construction. The honest, narrow floor a model has to clear.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1120 input and 273 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's vendor inquiries are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any text the model actually read this run. In a real deployment the vendor inquiry text would arrive from an external vendor's own email or portal message -- exactly the kind of externally-authored input fin-close's basis-document attack measures for a different artifact, and this kit's architecture treats it as trusted input with no verification step. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and two boundaries that should exist do not
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every vendor inquiry is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's inquiry would carry externally-authored text (see posture), and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it; two are negative results, not guarantees. Confirmed by reading the code, not by a run, on 2026-08-19 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a trace ever release a payment or change a remittance detail?
A drafted stage trace or reply could plausibly trigger a payment release or a remittance edit.
No code path does. src/payrun.py::check() and src/app.py's /api/check both return an answer only; neither writes to data/invoices.jsonl or any other store -- confirmed by reading every call site.
Can the drafted reply commit to a date beyond the invoice's own scheduled run?
A crafted inquiry or a model reply could invent an earlier or accelerated payment date.
src/prompt.py's SYSTEM states stated_date must be the real scheduled_date or remittance_date from the record, or null -- but this is a PROMPT-LEVEL rule; nothing in src/payrun.py or src/app.py cross-checks the model's stated_date against the record before rendering it. See Guardrails.fails.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/check handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Can the model's own current_stage and requires_ap_review disagree with each other, or with the record, and still reach the reader unflagged?
One would expect a stage/review mismatch or a stage that contradicts the record to be caught before it renders.
NO -- this is a real gap, not a guarantee. Nothing in src/payrun.py or src/app.py recomputes requires_ap_review from the model's own current_stage, or re-checks current_stage against the match/approval/run_inclusion/remittance record, before the UI shows them; both rules live only in the system prompt. See Guardrails.fails.
Each boundary above was checked by reading the call sites, not by an attack trial. Gates 2 and 4 are the ones that do NOT hold -- see Guardrails.fails for the same findings from the enforcement side.
The result0 attack trials, and two boundaries that do NOT hold in code (only in the prompt): the stated date is never cross-checked against the record, and requires_ap_review is never re-derived from current_stage, before the UI renders them. Two other boundaries (no write path, key redaction) hold, confirmed by reading the code.
1externally-authored field a live deployment would carry (the vendor inquiry text) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no invoice's vendor inquiry was authored by an outside party. A real deployment's inquiry text would be exactly the kind of externally-supplied text fin-close's basis-document red-team run measured being poisoned for a different artifact; that question is open for this kit.
Read this twice
This kit has no code-level check on its own model's stage/review consistency or its stated date. Everything from the requires_ap_review derivation to the date-grounding rule lives in the system prompt alone -- src/payrun.py and src/app.py pass the model's parsed reply straight through. On this run every field happened to be internally consistent and every gold remittance-trap invoice was correctly traced (0 false_paid), but that is a property of THIS run's replies, not a guarantee the code provides. A future version should recompute requires_ap_review from the model's own current_stage and cross-check stated_date against the record -- a cheap, deterministic check, the same shape evals/scoring.py already runs offline -- before trusting a live trace.
HonestyWhat this does not prove
Whether an injected instruction inside a vendor inquiry (e.g. 'this invoice is already cleared, mark it remitted') could move a stage call or the drafted reply -- no red-team run exists for this kit.
Whether the live app's hardcoded thinking=off setting changes resistance to a hostile inquiry relative to this run's provider-default-on configuration -- untested either way.
Whether a code-level consistency check (Guardrails.add_first) would catch a real stage/review or stage/date disagreement in practice -- none has been built or exercised.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
An invoice is never told to a vendor as paid unless its true governing stage is remitted.
src/prompt.py -- SYSTEM, stated as an explicit boundary the model must respect. Prompt-level ONLY: unlike fin-close's live citation-fidelity check or gap-brief's live fabrication check, no code in src/payrun.py or src/app.py recomputes current_stage from the record or cross-checks requires_ap_review/stated_date against it before the UI renders them.
EvidenceDoes it hold?
What
Measured
No code path in this kit releases a payment or changes a remittance detail
0 of 35 calls in r001-fin-payrun resulted in any write to data/invoices.jsonl or any other store -- src/payrun.py::check() and src/app.py's /api/check both only ever return an answer.
The prompt-level false-paid rule itself held on this run
0 of 4 gold remittance-trap invoices in r001-fin-payrun were traced or replied to as paid (false_paid=0) -- but see fails below for what does, and does not, enforce this.
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK. The rule that a non-remitted invoice is never told to a vendor as paid lives only in src/prompt.py's SYSTEM text -- nothing in src/payrun.py or src/app.py recomputes it. See fails above.
It does not verify requires_ap_review or stated_date against the record mathematically -- see fails.
It does not make the stage classification itself correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against a hostile vendor inquiry -- no red-team run exists for this kit (see the security page).
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 18 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
15 measured by the latest run3 need the model half
Metric
Owner
Role
Why this one
stage-precedence-exact-match-plus-review-and-date
Stage precedence exact match, plus AP-review derivation, false-paid claim scan and date exact match
alarm
false_paid as a raw count, never folded into stage_accuracy -- it is the expensive-direction error; the per-stage breakdown specifically, since a model that checks downstream fields first fails exactly on approval_exception and match_exception (see baseline_note); answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed — alarm on Any nonzero false_paid on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
35
different corpus — nothing is comparable
corpus.bytes
21,362
invoices edited — the count held, the bytes did not
split.count
35
the characters per invoice's assembled record block count moved — a different set was scored
split.size_p50
410
the median size of one characters per invoice's assembled record block moved
split.size_p95
513
the 95th-percentile size of one characters per invoice's assembled record block moved
dataset.rows
35
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
stage accuracy
not yet known
35 answered invoices
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat -- r001-fin-payrun ran once.
false paid
0 -- zero this run, this kit's headline guardrail metric
4 gold remittance-trap invoices
measured, one run
requires_ap_review accuracy
not yet known
35 invoices
100% this run. No repeat -- r001-fin-payrun ran once.
date accuracy
not yet known
35 invoices scored for stated date
100% this run. No repeat -- r001-fin-payrun ran once.
reasoning-token share of output
not yet known
35 calls
72.7% this run (6,960 of 9,575 output tokens), one run, one setting -- reasoning left at provider default, never compared against the live app's thinking-off setting. Narrative only, not a registered metric key -- the run record does not carry a reasoning-token-share field, only raw input/output totals.
measured, one run -- the per-stage breakdown of the overall stage-accuracy figure above; not yet known to be stable across a repeat.
invoices answered
not yet known
35 invoices
100% this run -- every invoice drew a stage classification, none silently dropped. No repeat -- r001-fin-payrun ran once.
latency
not yet known
35 calls
p50 2,629ms, p95 4,190ms on r001-fin-payrun -- reasoning left on. One recorded run, not a distribution.
input volume
0 -- fixed by the corpus and the prompt, not the model
35 calls
39,200 input tokens on r001-fin-payrun. Any movement means the prompt or the corpus changed.
output volume
not yet known
35 calls
9,575 output tokens on r001-fin-payrun -- model-specific, and includes whatever reasoning the provider chose to spend.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-fin-payrun 2026-08-19
approval exception accuracy, %
100.0
awaiting run inclusion accuracy, %
100.0
date accuracy, %
100.0
false paid
0
false paid rate, %
0.0
in scheduled run accuracy, %
100.0
input tokens, whole run
39200
model latency p50 ms
2629.00
model latency p95 ms
4190.00
match exception accuracy, %
100.0
output tokens, whole run
9575
remitted accuracy, %
100.0
review accuracy, %
100.0
stage accuracy, %
100.0
stage answered, %
100.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 15 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether reasoning is left on or explicitly disabled
not measured -- this run left it on (provider default); the live app always disables it; the two have never been compared on this corpus
reasoning
src/app.py hardcodes thinking=THINKING_OFF; r001-fin-payrun's own top-level thinking field is null (provider default), while every one of its 35 per-invoice records carries a nonzero reasoning_tokens (50-925). See Business.not_good_enough.
the planted trap (a downstream field that looks complete while an earlier stage governs, TRAP_N=6 of 14 exception-stage invoices in tools/build_corpus.py)
stage accuracy on the 6 trap invoices: 0% (rule-order floor, all 6 misread) -> 100% (the fast tier) on the identical invoices; false_paid specifically: 100% (4 of 4, floor) -> 0% (fast tier)
measured
results/eval-b000-ruleorder.json vs results/eval-r001-fin-payrun.json, same 35 invoices, one variable (which classifier reads the record).
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
stage accuracy
nothing yet
false paid
any nonzero value, on any run
requires_ap_review accuracy
nothing yet
date accuracy
nothing yet
reasoning-token share of output
nothing yet
per-stage accuracy
any stage falling below 100% on a re-run
invoices answered
nothing yet
latency
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
nothing yet
NextThe three you would add first
A live, code-level consistency check: re-derive requires_ap_review from the model's own current_stage (current_stage in (match_exception, approval_exception)) and flag any disagreement before the UI renders it.Currently nothing catches a model that reports an internally-inconsistent stage/review pair -- the same 'verdict disagrees with its own extracted values' shape as fin-close's false-clean cases, and this kit has no check for it at all, live or offline, beyond the eval's own comparison to gold.
A live check that stated_date matches the record's real scheduled_date (for in_scheduled_run) or remittance_date (for remitted), and is null otherwise.The system prompt states the rule; nothing in code enforces it before a reader sees it. See fails.
A red-team run against the vendor inquiry text, the way fin-close attacked its basis document.In a real deployment the inquiry arrives from an external vendor this kit's architecture treats as trusted with no verification step -- unmeasured for this kit (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting in src/payrun.py.
What this cannot tell you
Whether the same 100% figures hold with reasoning explicitly disabled, matching the live app's own setting -- see Business.not_good_enough.
Whether a live, code-level consistency check (see add_first) would ever have caught a real disagreement -- this run's own model never produced one to test against.
Whether the prompt-only false-paid and date-grounding rules hold against a hostile or malformed vendor inquiry -- no red-team run exists for this kit.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is an empty file with a comment explaining that the emptiness is load-bearing. The whole tracing decision is three files: src/prompt.py, src/payrun.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
35 invoices, generated from a fixed seed, never fetched. Recomputing gold from the same generated 4-stage record (never the target_stage that seeded it) via derive_stage() is what keeps the internal-consistency check honest -- see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the five-stage taxonomy, the precedence rule and the answer schema are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. Passing thinking through is a kwarg on one call site, not a client-library upgrade.
evaluation
evals/scoring.py
eval harnesses
exact match over a five-value stage vocabulary plus a boolean and a keyword scan for a paid claim is a dict comprehension and a regex, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per invoice -- load, assemble, prompt, call, parse -- with no branching and no state carried between invoices. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A gating stage this kit's fixed five-value vocabulary doesn't name needs a hand-written STAGES/STAGE_MEANINGS entry plus new corpus template logic in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here.
No built-in output validation beyond src/prompt.py::parse()'s own tolerant-but-not-creative JSON parse -- a framework with a schema-and-consistency-check layer baked in might catch a stage/review disagreement before it renders; this kit's stdlib-only design did not build one (see Guardrails.fails).
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-fin-payrun on the fast tier, 2026-08-19. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
2,629 ms
not yet known
nothing yet
Model, p95
4,190 ms
not yet known
nothing yet
Input tokens
39,200
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
9,575
not yet known
nothing yet
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-19, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
invoice records and vendor inquiries
data/invoices.jsonl -- 35 invoices, generated once from a fixed seed by tools/build_corpus.py; nothing else knows what an invoice's 4-stage record is
one whole invoice record -- all four stages -- sent in the one call; src/app.py and src/payrun.py both read it by invoice_id, never modified after generation
gold trace
data/gold.jsonl -- 35 rows, computed by tools/build_corpus.py's derive_stage() by re-reading the same generated 4-stage record, never carried over from the target_stage that seeded it
never -- evals/scoring.py is pure code, no model, no key. src/payrun.py::load_gold()'s own docstring: NEVER read by check().
the key
.env -- never committed
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 35 invoices and their vendor inquiries -- byte-identically from a fixed seed (SEED=20260819) every time it is run, with an independent --verify pass re-deriving every gold row from the invoice records actually written to disk.
21,362 bytes, one file (data/invoices.jsonl), 35 invoices -- see Data.corpus. --verify confirmed 0 drift across all 35 rows this session. (lenses.Data.corpus, tools/build_corpus.py)
a real deployment's payment-run trail is continuous, not a fixed seed, and a trail with a genuinely incomplete step (rather than an unfavourable one) cannot be recognised as such by this kit at all (see Data.breaks_on) -- neither was measured
point tools/build_corpus.py at your own vendors and stage reasons, and every published accuracy/false_paid figure is void -- they are this corpus's own phrasing templates and trap design (see Data.bring_your_own_boundary), not a property of the model
model
The five-stage taxonomy (match_exception, approval_exception, awaiting_run_inclusion, in_scheduled_run, remitted), the read-in-order precedence rule and the derived requires_ap_review boolean are declared once in src/prompt.py and sent in full inside the system prompt on every call -- a FIXED rulebook that does not vary by invoice. One call per VENDOR INQUIRY carries the invoice's full 4-stage record alongside it, behind src/adapters/__init__.py, reasoning left at the provider's default (on) -- the configuration r001-fin-payrun ships, and a DIFFERENT configuration from the one src/app.py's live UI actually runs (thinking explicitly off).
The rulebook portion of the system prompt is 989 (proportional-share) of 1,120 input tokens on the published call (lenses.LLM.prompt_parts[0]), sent identically on all 35 calls in r001-fin-payrun -- see LLM.prompt_verbatim for the exact text. 35 of 35 invoices in r001-fin-payrun returned a reply that parsed cleanly (0 failures, finish_reason 'stop' on all 35), with the largest single call using 1,017 of the 3,000-token MAX_TOKENS budget (925 of it reasoning) -- real headroom, unlike fin-invval's own near-miss at a smaller ceiling. (lenses.Business.not_good_enough, results/eval-r001-fin-payrun.json, src/payrun.py)
a gating stage this kit's fixed five-value vocabulary doesn't name needs a hand-written STAGES/STAGE_MEANINGS entry plus new corpus template logic in tools/build_corpus.py -- never learned from data alone, and never measured here. Separately, a call needing meaningfully more reasoning than this run's worst case (925 tokens) was never measured against MAX_TOKENS=3000.
a different stage set (added, removed or reordered gating stages) invalidates the whole eval at once -- gold and grading in evals/scoring.py are both keyed to this exact five-value vocabulary and its precedence order. Verdicts are also per-model and this configuration ran once -- the live app's own thinking-off setting has never been scored against this corpus; see Business.not_good_enough.
labels
data/gold.jsonl, 35 rows -- current_stage, requires_ap_review and expected_date are all RECOMPUTED from the same generated 4-stage record the model reads by derive_stage(), never from the target_stage that seeded the generator, and independently re-derived by --verify against the invoice records actually written to disk.
your own payment-run trail: hand-label the gold, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-labelled gold has an error rate this kit has never measured
stage accuracy and false_paid figures over this set reflect ONE planted trap family (a downstream field that looks complete while an earlier stage governs) and ONE assumption (every gating system is directly queryable) -- a real payment run's other failure shapes (see Data.breaks_on) are untested
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
current_stage is remitted, or the reply otherwise claims payment happened or is imminent, but the invoice's match or approval record is not actually clean
nothing in code would catch this -- the rule is prompt-only (see Guardrails.fails); this run never produced one, but the code does not prevent it
re-read match.matched and approval.status yourself before trusting a remitted verdict or a reply that claims payment, until Guardrails.add_first's live check exists (lenses Guardrails.fails, src/payrun.py, src/app.py)
current_stage is approval_exception or match_exception but requires_ap_review is false, or the reverse
the model stated the derived boolean inconsistently with its own stage call -- REVIEW_STAGES in src/prompt.py defines the rule but nothing in code re-derives requires_ap_review from current_stage before rendering; this run never produced this disagreement (see Eval.scores) but the code does not guard against it
recompute requires_ap_review = current_stage in (match_exception, approval_exception) yourself before trusting the field as printed (lenses.Guardrails.fails, src/prompt.py)
No machine symptom — this failure leaves no trace in any output.
reasoning ('thinking') left at provider default for the registered run, while the live app hardcodes it off -- a reader who re-runs this kit's own app will see different latency and cost than this report publishes, and nobody has measured whether accuracy differs too. See Business.not_good_enough.
Concurrency and GPU sizing -- one serial call per invoice, nothing measured past 35. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- Cost prices every call at the peak, cache-miss rate for exactly this reason. Whether a real vendor's own inquiry phrasing or a real payment system's own trap shapes would reproduce these figures -- see Data.bring_your_own_boundary. Whether reasoning explicitly disabled (the live app's own setting) changes any figure on this page -- every scored run here left it at provider default. Whether a hostile or malformed vendor inquiry could move a stage call or the drafted reply -- no red-team run exists for this kit.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Stage precedence exact match, plus AP-review derivation, false-paid claim scan and date exact match
Tell a vendor exactly where their invoice is stuck
PresenterOpens the private repo. Visible to admins only.
In one lineStage precedence exact match, plus AP-review derivation, false-paid claim scan and date exact match
Does the model's current_stage match gold's precedence-derived stage? Does requires_ap_review match the stage-derived boolean? Does the reply, or the stage itself, falsely claim the invoice is paid when a gold remittance-trap invoice's true stage is still an open exception? Does stated_date match the one real date the record supports, or correctly state none?
$0.00per 1,000 invoices
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Invoice
INV-2024-1020
Vendor
Oakmere Print & Design
Vendor inquiry (verbatim)
Following up on INV-2024-1020 -- any update on payment timing?
Gold stage
approval_exception
Rule-order floor's stage
remitted
The model's stage
approval_exception
The model's drafted reply
We are reviewing INV-2024-1020 because it requires additional approval before payment can be scheduled. We will provide an update once Accounts Payable completes its review.
Scored as
correct -- the remittance-trap case the free rule-order floor fails (a false paid claim) and the fast tier does not
Grader
Verdict
Why
Stage precedence exact match, plus AP-review derivation, false-paid claim scan and date exact match
correct
INV-2024-1020: match clean, approval exception (amount exceeds PO tolerance), but run_inclusion AND remittance both show a real run_id/scheduled_date and remittance_date/reference -- the invoice was later pulled back from the run it was scheduled on and paid, so both downstream fields look complete while approval is what actually governs. Gold stage approval_exception (a remittance-field trap). The rule-order floor checks remittance first and answers remitted -- a false paid claim. The fast tier answered approval_exception, correctly withheld any date, and did not claim payment in its drafted reply.
The formulaWhat it computes
stage_accuracy = correct stages / answered invoices, over stage_answered=35. false_paid = current_stage=='remitted' OR reply matches a plain-language paid-claim regex, counted over exactly the gold remittance-trap denominator (4), never folded into stage_accuracy. review_accuracy and date_accuracy are their own exact-match rates.
The analysisWhat it actually did
Model
Result
the fast tier
100.0% stage accuracy · 1 more measured on this row
In operationWhat to monitor
Reference standard: tools/build_corpus.py's derive_stage(), which reads the same 4-stage record IN ORDER and computes the gold current_stage, requires_ap_review and expected_date -- re-derived independently via --verify against the invoice records actually written to disk, asserted zero drift.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own planted trap design, which comes from tools/build_corpus.py's fixed templates.
Watch these
false_paid as a raw count, never folded into stage_accuracy -- it is the expensive-direction error
the per-stage breakdown specifically, since a model that checks downstream fields first fails exactly on approval_exception and match_exception (see baseline_note)
answered vs asked -- 100% this run, but a run that returns nothing has not scored well on what it managed
Alarm on
Any nonzero false_paid on any run -- the guardrail this kit's own prompt states as a boundary, not a hint. Zero this run.
How tight can the band be? There is no tolerance band here -- every field is an exact match (stage, boolean, date string) or a plain keyword scan; the AP-tracing task has no continuous quantity to round.
Cadence: Re-run on any change to src/prompt.py, tools/build_corpus.py, or MAX_TOKENS in src/payrun.py -- the first changes what is asked, the second changes what is asked ABOUT, the third bounds how many verdicts can come back at all.
The decisionWhen to reach for it
Use it
The gold stage is DERIVED from the same generated 4-stage record the model reads -- true of every kit corpus, never true of a real deployment's own payment-run trail.
Do not use it
The truth is not known in advance -- the normal state of a real AP inbox, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning