Check a marketplace seller's application against their documents
Every new seller's application has to match their business registration, bank letter and ID, field by field. This app checks all seven fields and drafts a summary of what does not match for your analyst.
PresenterOpens the private repo. Visible to admins only.
For seller verificationCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
A verification team at an online marketplace, reviewing each new seller before they can start selling.
✕Today's manual process
1Open three documents for every application: the business registration, the bank letter and the owner's ID.
2Compare seven fields one at a time: name, address, tax ID, bank details and owner.
3Judge the seller's note and decide whether it really explains a difference.
4One missed bank number on a clean-looking application lets a fraudulent seller get paid.
Every field compared manually
✓With the app
1The application arrives with its three documents already side by side.
2Each field is marked: matches, does not match, or explained by the seller.
3A note counts only when it actually speaks to the difference it covers.
4Your analyst gets a summary and still makes every approve or deny decision.
Analysts review only what differs
See it work
One real seller application: what the app checks, step by step
Bellwater Bicycle Parts Inc. declares an address in Alabama, but its business registration shows one in Wisconsin.
Check a marketplace seller's application against their documentsReference appBuilt to be shaped to your process
5
1The seller's application declared fields beside what the business registration, bank letter and ID show.
2The address differs declared in Alabama; the business registration shows Wisconsin.
3The bank routing number matches the bank letter exactly, so it is not flagged.
4The seller's note says nothing about the address, so it explains nothing.
5The outcome one field out of seven is flagged for the analyst.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check a marketplace seller's application against their documents
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Cross-checking a seller's onboarding application means comparing all seven declared fields against three submitted documents, deciding whether a submission note genuinely explains any discrepancy it finds -- never softening a flag because the rest of the application looks clean -- for every application, every submission. A verification analyst reading a seller's onboarding application against its three submitted documents by hand, checking all seven declared-vs-document field pairs for a discrepancy, and deciding whether a submission note genuinely explains any of them -- for every application, every submission.
Audience
Marketplace trust-and-safety and verification staff who review seller onboarding applications, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual applications
The corpus is 36 applications, 0.03 MB (jsonl 1). A real seller onboarding application names a real business, a real owner and a real bank account -- this repo has never held one and is not able to. There is no public corpus of paired (application, submitted documents, adjudicated verification outcome) records, for the same reason no public corpus of KYC decisions or fraud cases exists: the interesting cases are exactly the ones nobody can publish. Every business, owner, address, tax ID and bank detail in this corpus is invented.
The corpus
The 36 applicationsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your applications. That is the whole change — there is no database to migrate.
One application, as the model receives itapplications.jsonl · 1 of 36
A match/mismatch/mismatch_explained flag for each of the seven declared-vs-document field pairs, and a short summary drafted for a verification analyst -- informational only, never an approve/deny/escalate decision.
And when it cannot
Calling a genuinely mismatched bank_routing_number anything other than mismatch on an application that looks clean everywhere else (missed_banking_mismatch), or reading an absent or unrelated submission_note as covering a discrepancy it never mentions (over_explained). Measured at 0 of 13 planted banking-trap applications and 0 of 25 unexplained-mismatch cells this run -- see not_good_enough for why zero on one synthetic run is not the same claim as zero on a real application queue.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Deciding whether a lone unexplained banking mismatch on an otherwise-clean application should be flagged — the fast tier, over the free note-presence floor 0 of 13 planted banking-trap applications missed this run, against the floor's 10 of 13 (76.9% miss rate) -- the floor treats any attached note as covering the mismatch, regardless of what it actually says.
Deciding whether a submission note genuinely explains a specific field's discrepancy — the fast tier's mismatch_explained calls, read against which field the note actually names 0 of 25 unexplained-mismatch cells over-explained this run, against the floor's 21 of 25 (84.0%) -- the floor cannot tell a note that explains THIS field from a note that merely exists.
What this kit is not — a set of field flags and a drafted summary for a verification analyst to review -- never an approve/deny/escalate decision src/onboard.py and src/app.py have no function anywhere that writes a verification_outcome; every flagged inconsistency stays in the output, none auto-cleared.
At a glanceHow the whole thing runs
100%field accuracy pct
3,139 msp50, end to end
$1.57per 1,000 applications · Google Gemini 3 Flash
Run once, for real, on 2026-08-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check a marketplace seller's application against their documents14 steps · 4 questions · run once, for real · 2026-08-20
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own businesses, owners, addresses and document values, or write your own data/applications.jsonl in the same shape -- src/onboard.py and src/prompt.py read an application by its own fields (declared, business_registration, bank_letter, id_document, submission_note) and do not care where the values came from. The measured 100% field accuracy and 0% missed_banking_mismatch / over_explained rates are THIS corpus's document-value phrasing (invented businesses, owners and addresses drawn from fixed name lists) and THIS corpus's explanation-note templates (seven fixed explain-note strings, one per field).Corpus lens →
When is this the wrong choice?
Avoid: The free note-presence floor for anything but applications with no note at all -- it is not a competitor, it is the honest floor a model has to clear. That is the case against the best-fitting scenario (“Deciding whether a lone unexplained banking mismatch on an otherwise-clean application should be flagged”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
This kit assumes every submitted verification document already arrives text-extractable. The business registration, bank letter and government ID blocks are clean fields, standing in for documents already OCR'd or otherwise extracted to text -- a low-quality photo submission that OCRs poorly has no clean text for this kit to read, and this kit has no way to notice that a document it was given is degraded rather than simply mismatched. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether the same 100%/0% figures hold on a second run, or on an application combining more than one of the five named patterns at once. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-20 — r001-mp-onboard. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- application APP-2024-0357's declared record, all three document blocks and its submission note, since data/applications.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Clicking Cross-check with no API_KEY returns a calm 200 explaining nothing was called. It cannot reproduce a flag, a score or a dollar figure without a key -- those are what results/eval-r001-mp-onboard.json already committed. The free floor (python -m evals.baseline) and the corpus self-check (python tools/build_corpus.py --verify) both run on a cold clone with no key at all.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
3,139 msp50, end to end
5,039 msp95
2 minclone to first result
What the clock covers. END-TO-END per application: one HTTP request carrying the declared record and all three document blocks, parsed into seven field flags and a drafted summary. No retrieval step -- loading an application from disk is pure code and costs no network time. Reasoning was left at the provider's default (on), so p50 and p95 both include a reasoning pass on every call; the p95 of 5,039 ms IS one specific application, APP-2024-0127, which also produced this run's largest reply (643 output tokens, 524 of them reasoning -- both the run's maximum). The fastest call took 1,479 ms.
Current processWhat it replaces
A verification analyst reading a seller's onboarding application against its three submitted documents by hand, checking all seven declared-vs-document field pairs for a discrepancy, and deciding whether a submission note genuinely explains any of them -- for every application, every submission.
Where it is not good enough
Four findings. FIRST, a live discrepancy between what this report measures and what a forker's own click actually gets: the registered run (r001-mp-onboard) left provider-side reasoning ('thinking') at its documented default -- ON -- while the shipped app's own endpoint (src/app.py's /api/check) hardcodes thinking=THINKING_OFF on every live call. Nobody reconciled the two before this run was registered as the measured fact. Reasoning consumed 66.5% of this run's total output-token budget (8,206 of 12,336 output tokens; 227.9 tokens/call average, up to 524 on one application) and is a real driver of both this kit's latency (p50 3,139 ms, p95 5,039 ms) and, since a provider bills reasoning tokens as completion tokens, its dollar cost -- the published cost and latency numbers on this report are therefore NOT what a reader gets from the shipped UI, and that configuration has not been separately measured. SECOND, this run's own 100% field accuracy and 0% missed_banking_mismatch / over_explained rates are measured on one 36-application run, not a distribution -- no repeat exists (see Eval.could_not_verify), and the trap denominators are exactly what a corpus of this scale honestly supports: 13 planted banking-trap applications and 25 unexplained-mismatch cells across all patterns, not a stable rate over a larger sample. A zero here is a reason to verify harder, not to trust more -- the corpus is synthetic, one model, one run, no red-team run was fired. What the run DOES establish is separability: the free note-presence floor over the identical corpus misses 10 of 13 planted banking traps (76.9%) and over-explains 21 of 25 unexplained cells (84.0%), so a zero here is a real pass rather than a task nothing could fail. THIRD, a structural limitation stated in the corpus's own documentation, not found by this run: this kit assumes every submitted verification document already arrives text-extractable -- the business registration, bank letter and government ID blocks are clean fields, standing in for documents already OCR'd or otherwise extracted to text. A low-quality photo submission that OCRs poorly has no clean text for this kit to read, and this kit has no way to notice that a document it was given is degraded rather than simply mismatched. FOURTH, every application carries at most one of five named patterns (clean, banking_trap, explained_single, escalated_nonbanking, denied_multi); a real application can carry discrepancies in more combinations than these five cover, and this run says nothing about how the model behaves on a combination it was never shown.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free note-presence baseline (evals/baseline.py, treats any non-empty submission_note as explaining every mismatch in the application) scores 91.7 pct field accuracy overall but misses 10 of 13 planted banking-trap applications — a 76.9 pct missed_banking_mismatch rate — exactly the trap this kit's model resolved correctly on all 13 planted instances. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
whether reasoning is enabled
src/onboard.py and src/app.py
check() takes a thinking kwarg; src/app.py's live /api/check hardcodes it OFF while the registered eval run (r001-mp-onboard) left it at provider default (on) -- a real, live discrepancy between what this report measures and what a forker's own click experiences today. See Business.not_good_enough.
the field-pair taxonomy
src/prompt.py
A declared field this kit's fixed seven-field FIELDS tuple does not carry, or a different document source per field, needs a rewrite of SYSTEM plus FIELD_SOURCES/FIELD_MEANINGS -- a prompt tweak alone will not teach a new pairing the corpus never plants.
the corpus
tools/build_corpus.py
Point it at your own business names, owners, addresses and document values, or write your own data/applications.jsonl in the same shape -- src/onboard.py and src/prompt.py read an application by its own fields (declared, business_registration, bank_letter, id_document, submission_note) and do not care where the values came from.
the trap fraction
tools/build_corpus.py (PATTERN_COUNTS)
Set by one module dict, not a single global -- a new corpus states its own pattern counts in the same five names.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 36 applications and their document blocks from a fixed seed (SEED=20260819) across five named patterns (8 clean, 13 banking_trap, 8 explained_single, 4 escalated_nonbanking, 3 denied_multi). Gold (the seven field flags and verification_outcome) is DERIVED from the same declared-plus-document record by derive_flags()/derive_outcome(), never from the pattern the builder intended -- an assert inside build_application() catches any drift between intent and derivation before the file is even written, and --verify independently re-checks every row against data/*.jsonl on disk.
the prompt
src/prompt.py
The seven field-pair taxonomy, the three-flag vocabulary and the output schema, declared once and read from here by build() and parse(). States both failure modes to the model explicitly: a single unexplained banking mismatch on an otherwise-clean application is not 'probably fine', and a note only explains a field it specifically names.
the AI layer
src/onboard.py
Loads one application, calls the model once with the declared record and all three documents, parses the reply into seven field flags and a summary. Never approves, denies, or escalates -- there is no function here or in src/app.py that writes a verification_outcome.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
spend control
src/budget.py
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
the app
src/app.py
Static files and two JSON endpoints (/api/application, /api/check) on the standard library, port 8791. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-mp-onboard) actually used; see Business.not_good_enough.
the scorer
evals/scoring.py
Field-level scoring, pure code: three-way exact match on every field pair, missed_banking_mismatch (the expensive-direction error, over exactly the 13 planted trap applications) and over_explained (the opposite-direction error, over all 25 unexplained-mismatch cells) reported as their own named figures, never folded into a generic accuracy bucket.
free baseline
evals/baseline.py
A note-presence rule, written the way a person free-texting a quick classifier would write one: if any submission_note is attached at all, treat it as explaining whatever field disagrees. Correct on every application with no note and every application with no mismatch; wrong exactly where a note is attached but does not actually name the field that disagrees -- most visibly on the corpus's planted banking trap.
Where it breaks at scale
Not on application count -- each application is independent and one call per application is linear. It breaks three other ways. FIRST, ON THE STRUCTURAL ASSUMPTION stated in the corpus's own documentation: this kit assumes every submitted document already arrives text-extractable, and has no way to notice that a document it was given is degraded (a low-quality photo submission that OCRs poorly) rather than simply mismatched -- a real deployment needs a manual fallback for exactly that case. SECOND, ON COMBINATION COVERAGE: every application in this corpus carries at most one of five named patterns; a real application can carry discrepancies in combinations these five patterns do not cover (two independently mismatched fields with two different, unrelated explanations, say), and this run says nothing about that case. THIRD, ON THE REASONING CEILING: MAX_TOKENS is fixed at 3,000 (src/onboard.py), sized generously up front rather than computed -- this run's worst observed call used 643 of the 3,000-token budget (524 of it reasoning), real headroom, but reasoning was left at provider default (on) rather than disabled and consumed 66.5% of the average output-token budget; a call needing meaningfully more reasoning than this run's worst case was never measured against the ceiling.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
The panel before any call: application APP-2024-0357 loaded, all seven declared-vs-document field pairs shown side by side with their document source labelled underneath, and the submission note ('Excited to start selling on the platform!') displayed below the fields. Every row shows a neutral '?' -- nothing has been flagged yet.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same application after pressing Cross-check with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error, and the result panel honestly reports 'No run yet' rather than a false 'all fields match' verdict. This is the honest failure state, not a staged one.failureOpen full size →
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
36applications
0.03 MiBjsonl 1
p50 754chars per characters per application's assembled document block
$0.00setup · 0.0s
How it is cutWhat one characters per application's assembled document block is
One application, its declared record and all three document blocks included whole -- no train/test split, every application checked once, in one call.
SetupWhat the setup figure measured
No index is built. src/app.py finds an application by scanning C.applications() for a matching application_id -- 36 rows, a linear scan, not a search structure -- and src/onboard.py sends that application's full record whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own applications
Point tools/build_corpus.py at your own businesses, owners, addresses and document values, or write your own data/applications.jsonl in the same shape -- src/onboard.py and src/prompt.py read an application by its own fields (declared, business_registration, bank_letter, id_document, submission_note) and do not care where the values came from. The three-flag vocabulary in src/prompt.py assumes exactly match, mismatch and mismatch_explained; a fourth state (a 'document illegible' flag, say) is a prompt.py AND tools/build_corpus.py change, not a data-only change.
⚠︎ And what stops being true when you do: The measured 100% field accuracy and 0% missed_banking_mismatch / over_explained rates are THIS corpus's document-value phrasing (invented businesses, owners and addresses drawn from fixed name lists) and THIS corpus's explanation-note templates (seven fixed explain-note strings, one per field). A real seller's own document formatting, and a real verification team's own reasons a field might genuinely disagree, are both unmeasured by this run.
What breaks it
This kit assumes every submitted verification document already arrives text-extractable. The business registration, bank letter and government ID blocks are clean fields, standing in for documents already OCR'd or otherwise extracted to text -- a low-quality photo submission that OCRs poorly has no clean text for this kit to read, and this kit has no way to notice that a document it was given is degraded rather than simply mismatched.
Every application carries at most one of five named patterns. A real application can carry discrepancies in more combinations than clean, banking_trap, explained_single, escalated_nonbanking and denied_multi cover -- two independently mismatched fields with two different explanations, say -- and this run says nothing about that case.
The trap denominator is small -- 13 banking-trap applications and 25 unexplained-mismatch cells across all patterns. That is what a corpus of this scale supports honestly, not a stable rate.
Both the corpus's own gold labels and its note text are drawn from a small, fixed set of templates -- data/SOURCES.md documents seven explain-note templates (one per field) and ten generic filler notes, and derive_flags() itself resolves 'explained' by a fixed keyword-phrase match against those same templates. The model reads the prose directly and was never shown the keyword list, but whether its judgment would generalize to a genuinely varied, freely-worded explanation outside this template set is unmeasured by this run.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,369
887
application
757
199
Total
1,086
This is the cost lesson as arithmetic: of the 1,086 tokens assembled, 887 are instructions — 82% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged by importing src.prompt directly and calling prompt.build(app) for APP-2024-0357 (not retyped) -- the literal system and user content src/onboard.py::check() sends. Per-part token counts are a proportional estimate over character share (system 3,369 / application 757 chars, 4,126 total) applied to this call's own total of 1,086 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split. The 'application' part covers only the declared-plus-document block src/prompt.py::_doc_block() builds -- the surrounding 'Cross-check the application below field by field...' instruction text is part of the user message actually sent (see prompt_verbatim) but is not a separately named segment in the kit's own code.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You extract and cross-check a third-party seller's onboarding application. You are given the applicant's declared business identity, banking details and owner identity; the business registration document; the bank verification letter; the government ID document; and the applicant's own submission note (which may be empty). All three documents are already extracted to clean text -- you are not reading images.
YOU NEVER APPROVE, DENY, OR ESCALATE THIS APPLICATION, AND YOU NEVER AUTO-CLEAR A FLAGGED INCONSISTENCY, EVEN ONE THAT LOOKS MINOR. That decision belongs to a verification analyst. Your only job is to extract each declared field, compare it against the matching document field, and flag every inconsistency -- every mismatch and every mismatch_explained stays in your output; there is no field you drop because it seems small.
THE SEVEN FIELD PAIRS, EACH FLAGGED INDEPENDENTLY
business_name declared business_name against the business registration document's legal_name
business_address declared business_address against the business registration document's registered_address
tax_id declared tax_id against the business registration document's own tax_id
bank_account_name declared bank_account_name against the bank letter's account_holder_name
bank_routing_number declared bank_routing_number against the bank letter's routing_number
owner_name declared owner_name against the government ID document's full_name
owner_address declared owner_address against the government ID document's address
FOR EACH FIELD, EXACTLY ONE OF THREE FLAGS
match the declared value and the document value agree
mismatch they disagree, and nothing in submission_note explains why
mismatch_explained they disagree, AND submission_note specifically and genuinely accounts for THAT EXACT field's discrepancy -- not a generic note, not a note about a different field
⚠︎ A SINGLE MISMATCH ON AN OTHERWISE-CLEAN APPLICATION IS NOT A REASON TO SOFTEN IT. An application where six of seven fields match cleanly and only bank_routing_number disagrees is not "probably fine" -- a mismatched banking detail on an application that looks clean everywhere else is the exact pattern a fraudulent submission produces. Flag it exactly as you would flag it on a messier application; do not let five or six clean fields talk you out of the one that isn't.
⚠︎ NEVER CALL A FIELD mismatch_explained UNLESS submission_note SPECIFICALLY NAMES THAT FIELD'S DISCREPANCY. An empty note, a generic note ("please process quickly"), or a note that explains a different field does not explain this one -- when in doubt, the correct flag is mismatch, not mismatch_explained. A genuinely explained discrepancy IS a real, legitimate outcome -- "recently changed business address, registration update pending" does explain a business_address mismatch -- but the note has to actually say so for this field, not merely exist.
Return a JSON object with exactly these keys: business_name, business_address, tax_id, bank_account_name, bank_routing_number, owner_name, owner_address (each one of: match, mismatch, mismatch_explained), and summary (one or two sentences drafted for the verification analyst, naming which fields you flagged and why -- never a decision to approve, deny, or escalate the application).
Cross-check the application below field by field and draft a summary for the verification analyst.
Return a JSON object with exactly eight keys: business_name, business_address, tax_id, bank_account_name, bank_routing_number, owner_name, owner_address (each one of: match, mismatch, mismatch_explained), and "summary" (the drafted note, one or two sentences).
APPLICATION APP-2024-0357
DECLARED ON THE APPLICATION FORM
business_name: Bellwater Bicycle Parts Inc.
business_address: 2667 Juniper Ln., Northport, AL 35473
tax_id: 90-2443853
bank_account_name: Bellwater Bicycle Parts Inc.
bank_routing_number: 155567574
owner_name: Yuki Norgaard
owner_address: 8070 Hazel St., Concord, NH 03301
BUSINESS REGISTRATION DOCUMENT
legal_name: Bellwater Bicycle Parts Inc.
registered_address: 8623 Poplar Ct., Ashland, WI 54806
tax_id: 90-2443853
BANK VERIFICATION LETTER
account_holder_name: Bellwater Bicycle Parts Inc.
routing_number: 155567574
GOVERNMENT ID DOCUMENT
full_name: Yuki Norgaard
address: 8070 Hazel St., Concord, NH 03301
SUBMISSION NOTE: Excited to start selling on the platform!
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"business_name": "match",
"business_address": "mismatch",
"tax_id": "match",
"bank_account_name": "match",
"bank_routing_number": "match",
"owner_name": "match",
"owner_address": "match",
"summary": "Only business_address is flagged: the declared address (2667 Juniper Ln., Northport, AL 35473) differs from the registered address on the business registration document (8623 Poplar Ct., Ashland, WI 54806). The submission note does not explain this discrepancy."
}
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check a marketplace seller's application against their documents — 36 applications. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: a three-way exact match on every field pair (match/mismatch/mismatch_explained), missed_banking_mismatch (a wrong bank_routing_number flag on one of the 13 planted trap applications) reported as its own named figure, and over_explained (a mismatch_explained flag on any of the 25 unexplained-mismatch cells across all patterns) reported separately again. No LLM judge anywhere in the path. The same function scores both evals/baseline.py's free floor and evals/run.py's real run.
36applications
36source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED252 / 252field accuracy pct — field-pair flags answeredDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED252 / 252field answered pct — field-pair flags askedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 13missed banking mismatch rate pct — planted banking-trap applicationsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 25over explained rate pct — unexplained-mismatch cells across all patternsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer is pure code, exact match against a mechanically-derived gold set: every application is tagged with its seven field flags and its verification_outcome at generation time by derive_flags()/derive_outcome(), with an assert inside build_application() catching any drift between the pattern the builder intended and what derive_flags() actually computes before the file is even written -- and, independently, tools/build_corpus.py --verify re-derives every row from data/applications.jsonl on disk and asserts zero drift, confirmed this session (36 applications checked, 0 drift; 13 trap applications, all trap_field=bank_routing_number, all mismatch unexplained, all denied).
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One application
1,000 applications
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.001570
$1.57
35%
Same work, 1× the bill
The same applications, the same tokens — only the rate card changed. And on that card about 35% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether reasoning ('thinking') is left on or explicitly disabled for this model -- the app's own /api/check always disables it; the registered run left it at provider default. The two configurations have not been priced against each other on this corpus.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page. Reasoning was also left at provider default (on) for this run -- see Business.not_good_enough -- so even the projected figure prices tokens the shipped app's own reasoning-off configuration would not spend.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. Field accuracy, missed_banking_mismatch and over_explained are all pure code over the committed run file, so a forker re-scores this run -- or the free floor -- for $0.00 and needs no key.
The gradersOne way to grade, and why it is the only one
The floor is correct on every application with no note at all and every application with no mismatch at all. It goes wrong exactly where a note is attached but does not actually address the field that disagrees -- most visibly on the corpus's planted banking trap, where 10 of the 13 trap applications carry a generic, unrelated note (drawn from GENERIC_NOTES) that this baseline reads as an explanation it isn't: those 10 are misread as mismatch_explained (76.9% missed_banking_mismatch; the other 3 trap applications carry no note at all, so the floor happens to get them right too). 21 of the 25 unexplained-mismatch cells across all patterns are over-explained the same way (84.0%). Its overall field accuracy still reaches 91.7% (231 of 252) because most of the 252 field cells are simple matches where both the floor and the model agree. The fast tier's 0% on both expensive-direction metrics on the identical 36 applications is the real, measured gap a model closes that this floor cannot.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Field-pair exact match, plus banking-mismatch and over-explanation checks Does each of the seven field pairs' match/mismatch/mismatch_explained call match gold? Does the model avoid calling a genuinely mismatched bank_routing_number anything other than mismatch on the 13 applications where it is the ONLY field that disagrees (missed_banking_mismatch)? Does it avoid calling a genuinely unexplained mismatch mismatch_explained on any of the 25 cells across all patterns where gold is mismatch (over_explained)?
$0.00
no
yes
the fast tier 100.0% field accuracy
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, and on both axes this corpus was built to test. The free floor and the fast tier are scored by the identical function over the identical 36 applications, and they separate cleanly: 91.7% vs 100% field accuracy, 10 vs 0 missed banking mismatches (of 13), and 21 vs 0 over-explained cells (of 25) -- exactly the corpus's planted trap counts. A grader that could not tell the two apart would not produce a gap that lines up this precisely with the corpus's own documented trap design.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Deciding whether a lone unexplained banking mismatch on an otherwise-clean application should be flagged
the fast tier, over the free note-presence floor
0 of 13 planted banking-trap applications missed this run, against the floor's 10 of 13 (76.9% miss rate) -- the floor treats any attached note as covering the mismatch, regardless of what it actually says.
the free note-presence floor for anything but applications with no note at all -- it is not a competitor, it is the honest floor a model has to clear.
Deciding whether a submission note genuinely explains a specific field's discrepancy
the fast tier's mismatch_explained calls, read against which field the note actually names
0 of 25 unexplained-mismatch cells over-explained this run, against the floor's 21 of 25 (84.0%) -- the floor cannot tell a note that explains THIS field from a note that merely exists.
trusting a mismatch_explained call on a note phrased very differently from this corpus's seven fixed explain-note templates -- see Data.breaks_on; that generalization is unmeasured.
What this kit is not
a set of field flags and a drafted summary for a verification analyst to review -- never an approve/deny/escalate decision
src/onboard.py and src/app.py have no function anywhere that writes a verification_outcome; every flagged inconsistency stays in the output, none auto-cleared.
wiring this straight into an auto-approve/auto-deny pipeline with no analyst review step for any application carrying a mismatch flag.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
clean
All seven fields match cleanly, no discrepancy anywhere
8
APP-2024-0282: all seven fields match. Gold and model both: All seven declared fields match their corresponding supporting documents. No inconsistencies were identified for the verification analyst to review.
banking_trap
Every field matches except bank_routing_number, unexplained -- the planted trap
13
APP-2024-0662: gold flags bank_routing_number=mismatch. Gold and model both: All declared fields match their supporting documents except bank_routing_number: the application states 194350611 while the bank verification letter lists 208468486. No submission…
explained_single
One field mismatches and submission_note specifically names it
8
APP-2024-1145: gold flags business_address=mismatch_explained. Gold and model both: Only the business_address is flagged: the declared address differs from the registration document's registered_address, and the submission note explains that the address…
escalated_nonbanking
One non-banking field mismatches, unexplained
4
APP-2024-0357: gold flags business_address=mismatch. Gold and model both: Only business_address is flagged: the declared address (2667 Juniper Ln., Northport, AL 35473) differs from the registered address on the business registration document (8623 Poplar…
denied_multi
bank_account_name plus one or two other fields mismatch, unexplained
3
APP-2024-0919: gold flags tax_id=mismatch, bank_account_name=mismatch, owner_address=mismatch. Gold and model both: Flagged three inconsistencies: the declared tax_id (99-6773165) does not match the registration document (12-9884055); the declared…
What we could NOT verify
Whether the same 100%/0% figures hold on a second run, or on an application combining more than one of the five named patterns at once.
Whether a live, code-level consistency check (see Guardrails.add_first) would ever have caught a real disagreement -- this run's own model never produced one to test against.
Whether the prompt-only banking-trap and over-explaining rules hold against a submission_note crafted to mimic this kit's own explanation phrasing -- no red-team run exists for this kit.
Whether the same figures hold with reasoning explicitly disabled, matching the live app's own setting -- see Business.not_good_enough.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
1,083.5
342.67
3,139 ms
$0.001570
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The note-presence baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 39,006 input / 12,336 output tokens priced at Google Gemini 3 Flash's published rate -- the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The provider's reasoning pass, left at its default (on) rather than disabled for this run -- it consumed 66.5% of the average output-token budget (227.9 of 342.7 tokens/call) and is already priced into cost_per_query_usd, since a provider bills reasoning tokens as completion tokens. src/app.py's live UI hardcodes thinking off, a different, unmeasured configuration -- see Business.not_good_enough.
The declared record plus all three document blocks (business registration, bank letter, government ID) -- fixed size per call by construction (input tokens ranged 1,070-1,100 across all 36 calls, essentially flat), the floor every call pays regardless of how many fields end up flagged.
The fixed system prompt (3,369 characters -- the seven field-pair taxonomy, the three-flag vocabulary and the output schema) is sent in full on every call: 81.7% of the system+application character total on the published call.
Your volumeWhat it costs at your volume
Linear in applications: each call is independent and self-contained, with no shared context or retrieval step to amortise. This run's 36 applications cost about $0.0565 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $0.565 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volumeseller onboarding applications cross-checked per day/week -- this run checked 36 in one pass
What it replacesa verification analyst reading a seller's application against its three submitted documents by hand, checking all seven declared-vs-document field pairs and deciding whether a submission note genuinely explains any discrepancy it finds
Time saved per itemnot measured here -- depends on how long a manual document cross-check takes at the reader's own marketplace
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run, so this page prices one model, not a trade-off; see Eval.could_not_verify for what a second run would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
39,006input tokens · this run
12,336output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 36 applications cross-checked, 252 field-pair flags scored by pure code. This run (r001-mp-onboard) answered 252 of 252 field flags -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.023
$0.023
$0.63
2026-09-12
gemini-3-flash
Google
$0.057
$0.057
$1.57
2026-09-18
gemini-3-8-flash
Google
$0.076
$0.076
$2.10
2026-09-18
claude-haiku-4-5
Anthropic
$0.101
$0.101
$2.80
2026-09-12
llama-5
Meta
$0.101
$0.101
$2.81
2026-09-18
grok-4-5
xAI
$0.152
$0.152
$4.22
2026-09-18
grok-4-6
xAI
$0.152
$0.152
$4.22
2026-09-18
claude-sonnet-5
Anthropic
$0.201
$0.201
$5.59
2026-09-12
gemini-3-1-pro
Google
$0.226
$0.226
$6.28
2026-09-18
gpt-5-6-terra
OpenAI
$0.226
$0.226
$6.28
2026-09-12
gpt-5-6-sol
OpenAI
$0.403
$0.403
$11.19
2026-09-12
claude-opus-4-8
Anthropic
$0.503
$0.503
$13.98
2026-09-12
claude-opus-5
Anthropic
$0.503
$0.503
$13.98
2026-09-12
claude-fable-5
Anthropic
$1.007
$1.007
$27.97
2026-09-18
claude-fable-5-1
Anthropic
$1.007
$1.007
$27.97
2026-09-18
gpt-6-astra
OpenAI
$1.007
$1.007
$27.97
2026-09-17
Read this against the numbers above
REASONING WAS NOT DISABLED FOR THIS WORKLOAD. Every row below prices THIS run's own token counts, which include reasoning tokens the live app's own /api/check never generates (it hardcodes thinking off). A forker's real bill on other models depends on whether that model has an equivalent reasoning toggle and whether it is left on.
NO QUALITY IS IMPLIED. Only the model that produced Eval.scores has been scored against this corpus -- every row here is a price, not a recommendation.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 36 applications and their document blocks from a fixed seed (SEED=20260819) across five named patterns (8 clean, 13 banking_trap, 8 explained_single, 4 escalated_nonbanking, 3 denied_multi). Gold (the seven field flags and verification_outcome) is DERIVED from the same declared-plus-document record by derive_flags()/derive_outcome(), never from the pattern the builder intended -- an assert inside build_application() catches any drift between intent and derivation before the file is even written, and --verify independently re-checks every row against data/*.jsonl on disk.
You change it to: Point it at your own business names, owners, addresses and document values, or write your own data/applications.jsonl in the same shape -- src/onboard.py and src/prompt.py read an application by its own fields (declared, business_registration, bank_letter, id_document, submission_note) and do not care where the values came from.
tools/build_corpus.py
# Generate the seller onboarding applications this kit checks against, from a fixed seed.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
PATTERN_COUNTS = {
TRAP_N = PATTERN_COUNTS["banking_trap"]
BANKING_FIELDS = ("bank_account_name", "bank_routing_number")
KEYWORDS = {
EXPLAIN_NOTES = {
GENERIC_NOTES = (
src/prompt.pythe prompt — a swap seam
The seven field-pair taxonomy, the three-flag vocabulary and the output schema, declared once and read from here by build() and parse(). States both failure modes to the model explicitly: a single unexplained banking mismatch on an otherwise-clean application is not 'probably fine', and a note only explains a field it specifically names.
You change it to: A declared field this kit's fixed seven-field FIELDS tuple does not carry, or a different document source per field, needs a rewrite of SYSTEM plus FIELD_SOURCES/FIELD_MEANINGS -- a prompt tweak alone will not teach a new pairing the corpus never plants.
src/prompt.py
# Assemble the one prompt this kit sends, and parse the one reply it gets back.
FIELDS = (
FLAGS = ("match", "mismatch", "mismatch_explained")
FIELD_SOURCES = {
FIELD_MEANINGS = {
FLAG_MEANINGS = {
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _doc_block(app):
src/onboard.pythe AI layer
Loads one application, calls the model once with the declared record and all three documents, parses the reply into seven field flags and a summary. Never approves, denies, or escalates -- there is no function here or in src/app.py that writes a verification_outcome.
src/onboard.py
# Cross-check one seller onboarding application against its submitted documents, and flag every
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
APPLICATIONS = os.path.join(HERE, "data", "applications.jsonl")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
MAX_TOKENS = 3000
def applications():
def load_gold():
def check(cfg, app, complete=None, thinking=None, prompt=P.DEFAULT_PROMPT):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pyspend control
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
src/app.pythe app
Static files and two JSON endpoints (/api/application, /api/check) on the standard library, port 8791. /api/check hardcodes thinking=THINKING_OFF on every live call -- a different reasoning setting from the one the registered eval run (r001-mp-onboard) actually used; see Business.not_good_enough.
src/app.py
# The minimal local UI. Standard library only — python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8791"))
class H(BaseHTTPRequestHandler):
def main():
evals/scoring.pythe scorer
Field-level scoring, pure code: three-way exact match on every field pair, missed_banking_mismatch (the expensive-direction error, over exactly the 13 planted trap applications) and over_explained (the opposite-direction error, over all 25 unexplained-mismatch cells) reported as their own named figures, never folded into a generic accuracy bucket.
evals/scoring.py
# Score a set of predicted applications against gold. Pure code, shared by evals/baseline.py and
FIELDS = ("business_name", "business_address", "tax_id", "bank_account_name",
FLAGS = ("match", "mismatch", "mismatch_explained")
def score(records, gold):
evals/baseline.pyfree baseline
A note-presence rule, written the way a person free-texting a quick classifier would write one: if any submission_note is attached at all, treat it as explaining whatever field disagrees. Correct on every application with no note and every application with no mismatch; wrong exactly where a note is attached but does not actually name the field that disagrees -- most visibly on the corpus's planted banking trap.
evals/baseline.py
# What a simple, dumb rule catches, over the same corpus. Free. No key, no dependency, no model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def classify(app):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 36 applications and their document blocks from a fixed seed (SEED=20260819) across five named patterns (8 clean, 13 banking_trap, 8 explained_single, 4 escalated_nonbanking, 3 denied_multi). Gold (the seven field flags and verification_outcome) is DERIVED from the same declared-plus-document record by derive_flags()/derive_outcome(), never from the pattern the builder intended -- an assert inside build_application() catches any drift between intent and derivation before the file is even written, and --verify independently re-checks every row against data/*.jsonl on disk. A swap seam.
src/prompt.pyThe seven field-pair taxonomy, the three-flag vocabulary and the output schema, declared once and read from here by build() and parse(). States both failure modes to the model explicitly: a single unexplained banking mismatch on an otherwise-clean application is not 'probably fine', and a note only explains a field it specifically names. A swap seam.
src/onboard.pyLoads one application, calls the model once with the declared record and all three documents, parses the reply into seven field flags and a summary. Never approves, denies, or escalates -- there is no function here or in src/app.py that writes a verification_outcome.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Passes a thinking kwarg through to the provider only when the caller supplies one, and reports token_details.reasoning_tokens when the provider returns it -- the field the reasoning-token finding in Business.not_good_enough is measured from. A swap seam.
src/budget.pyAn append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
evals/scoring.pyField-level scoring, pure code: three-way exact match on every field pair, missed_banking_mismatch (the expensive-direction error, over exactly the 13 planted trap applications) and over_explained (the opposite-direction error, over all 25 unexplained-mismatch cells) reported as their own named figures, never folded into a generic accuracy bucket.
evals/baseline.pyA note-presence rule, written the way a person free-texting a quick classifier would write one: if any submission_note is attached at all, treat it as explaining whatever field disagrees. Correct on every application with no note and every application with no mismatch; wrong exactly where a note is attached but does not actually name the field that disagrees -- most visibly on the corpus's planted banking trap.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1083 input and 342 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's applications are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any field on it: not a declared value, a document, or the submission note. In a real deployment the declared fields, the three document blocks and the submission_note would all arrive from an external seller's own onboarding form and uploaded documents -- exactly the kind of externally-authored input this kit's architecture treats as trusted, with no verification step. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and two boundaries that should exist do not
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every declared value, document field and submission_note is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's declared record and documents are exactly the kind of externally-supplied text a seller controls, and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it; two are negative results, not guarantees. Confirmed by reading the code, not by a run, on 2026-08-20 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a cross-check ever approve, deny, or escalate an application?
A drafted set of field flags could plausibly be read by a downstream system as a decision.
No code path does. src/onboard.py::check() and src/app.py's /api/check both return an answer only; neither writes a verification_outcome anywhere -- confirmed by reading every call site.
Can a flagged inconsistency be auto-cleared because the rest of the application looks clean?
A reader (model or code) could soften or drop a flag because six of seven fields already look fine.
src/prompt.py's SYSTEM states the opposite rule in as many words -- but this is a PROMPT-LEVEL rule; nothing in src/onboard.py or src/app.py cross-checks the model's own seven flags against the declared/document values before rendering them. See Guardrails.fails.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/check handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Can the model's own seven flags disagree with its own drafted summary, and still reach the reader unflagged?
One would expect a flag/summary mismatch (e.g. bank_routing_number flagged mismatch but the summary text claims everything matches) to be caught before it renders.
NO -- this is a real gap, not a guarantee. Nothing in src/onboard.py or src/app.py cross-checks the drafted summary text against the seven parsed flags, or the flags against the declared/document values, before the UI shows them; the only check is src/prompt.py::parse()'s tolerant JSON parse. See Guardrails.fails.
Each boundary above was checked by reading the call sites, not by an attack trial. Gates 2 and 4 are the ones that do NOT hold -- see Guardrails.fails for the same findings from the enforcement side.
The result0 attack trials, and two boundaries that do NOT hold in code (only in the prompt): neither the model's own flags nor its drafted summary are cross-checked against the record before the UI renders them. Two other boundaries (no decision-write path, key redaction) hold, confirmed by reading the code.
4externally-authored field types a live deployment would carry (the declared record and all three document blocks) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no application's declared fields, document values or submission_note were authored by an outside party. A real deployment's declared record and documents would be exactly the kind of externally-supplied text a seller controls; whether an application authored to read as clean while carrying a crafted submission_note could talk the model out of a flag is unmeasured for this kit.
Read this twice
This kit has no code-level check on its own model's output. Nothing verifies that the seven flags are consistent with the declared/document values, or that the drafted summary agrees with the flags, before the UI renders them. On this run every field happened to be correct and every summary happened to match its own flags (0 missed_banking_mismatch, 0 over_explained), but that is a property of THIS run's replies, not a guarantee the code provides. A future version should recompute each field's match/mismatch call from the declared and document values directly and cross-check the summary text against the flags -- a cheap, deterministic check, the same shape evals/scoring.py already runs offline -- before trusting a live cross-check.
HonestyWhat this does not prove
Whether a submission_note crafted to mimic this kit's own explanation phrasing, for a field it does not actually address, could talk the model into a wrong mismatch_explained call -- no red-team run exists for this kit.
Whether the live app's hardcoded thinking=off setting changes resistance to a crafted application relative to this run's provider-default-on configuration -- untested either way.
Whether a code-level consistency check (Guardrails.add_first) would catch a real flag/summary or flag/value disagreement in practice -- none has been built or exercised.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
A single unexplained mismatch is never softened because the rest of the application looks clean, and a field is never called mismatch_explained unless submission_note specifically names that field's discrepancy.
src/prompt.py -- SYSTEM, stated as two explicit boundaries the model must respect. Prompt-level ONLY: no code in src/onboard.py or src/app.py recomputes a field's flag from the declared/document values, or cross-checks the drafted summary against the flags, before the UI renders them.
EvidenceDoes it hold?
What
Measured
No code path in this kit writes a verification_outcome or otherwise approves, denies, or escalates an application
0 of 36 calls in r001-mp-onboard resulted in any write or any outbound action -- src/onboard.py::check() and src/app.py's /api/check both only ever return an answer.
The prompt-level 'do not soften a lone banking mismatch' rule held on this run
0 of 13 planted banking-trap applications had their bank_routing_number flag softened to anything other than mismatch (missed_banking_mismatch=0) -- but see fails for what does, and does not, enforce this.
The prompt-level 'note must specifically name the field' rule held on this run
0 of 25 unexplained-mismatch cells across all patterns were called mismatch_explained on the strength of an absent or unrelated note (over_explained=0).
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK. Both rules live only in src/prompt.py's SYSTEM text -- nothing in src/onboard.py or src/app.py recomputes them. See fails.
It does not verify a flag against the declared/document values mathematically -- it trusts the model's own comparison.
It does not make the cross-check itself correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against a crafted submission_note -- no red-team run exists for this kit (see the security page).
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 15 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
17 measured by the latest run-2 need the model half
Field-pair exact match, plus banking-mismatch and over-explanation checks
alarm
missed_banking_mismatch as a raw count, never folded into field_accuracy -- it is the expensive-direction error this kit's trap exists to catch; over_explained specifically on applications with ANY submission_note, since a reader that treats note-presence as explanation fails exactly there (see baseline_note); answered vs asked -- 100% this run, but a truncated reply answers fewer fields while every field it DID answer stays correct, the same truncation-artefact shape sibling kits found at a smaller MAX_TOKENS — alarm on Any nonzero missed_banking_mismatch or over_explained, on any run. Both zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
36
different corpus — nothing is comparable
corpus.bytes
29,093
applications edited — the count held, the bytes did not
split.count
36
the characters per application's assembled document block count moved — a different set was scored
split.size_p50
754
the median size of one characters per application's assembled document block moved
split.size_p95
821
the 95th-percentile size of one characters per application's assembled document block moved
dataset.rows
36
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
field accuracy
not yet known
252 answered field-pair cells
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat -- r001-mp-onboard ran once.
field answered
not yet known
252 field-pair cells
100% this run -- every field-pair cell drew a flag, none silently dropped. No repeat -- r001-mp-onboard ran once.
missed banking mismatch
0 -- zero this run, this kit's headline guardrail metric
13 planted banking-trap applications
measured, one run
over-explained
0 -- zero this run, this kit's other headline guardrail metric
25 unexplained-mismatch cells across all patterns
measured, one run
per-field accuracy
100 -- 100% on all seven fields this run
36 gold cells per field
measured, one run -- the per-field breakdown of the overall field-accuracy figure above; not yet known to be stable across a repeat.
reasoning-token share of output
not yet known
36 calls
66.5% this run (8,206 of 12,336 output tokens), one run, one setting -- reasoning left at provider default, never compared against the live app's thinking-off setting. Narrative only, not a registered metric key -- the run record does not carry a reasoning-token-share field, only raw input/output totals.
latency
not yet known
36 calls
p50 3,139ms, p95 5,039ms on r001-mp-onboard -- reasoning left on. One recorded run, not a distribution.
input volume
0 -- fixed by the corpus and the prompt, not the model
36 calls
39,006 input tokens on r001-mp-onboard. Any movement means the prompt or the corpus changed.
output volume
not yet known
36 calls
12,336 output tokens on r001-mp-onboard -- model-specific, and includes whatever reasoning the provider chose to spend.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-mp-onboard 2026-08-20
bank account name accuracy, %
100.0
bank routing number accuracy, %
100.0
business address accuracy, %
100.0
business name accuracy, %
100.0
field accuracy, %
100.0
field answered, %
100.0
input tokens, whole run
39006
model latency p50 ms
3139.00
model latency p95 ms
5039.00
missed banking mismatch
0
missed banking mismatch rate, %
0.0
output tokens, whole run
12336
over explained
0
over explained rate, %
0.0
owner address accuracy, %
100.0
owner name accuracy, %
100.0
tax id accuracy, %
100.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 17 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether reasoning is left on or explicitly disabled
not measured -- this run left it on (provider default); the live app always disables it; the two have never been compared on this corpus
reasoning
src/app.py hardcodes thinking=THINKING_OFF; r001-mp-onboard's own top-level thinking field is null (provider default), while every one of its 36 per-application records carries a nonzero reasoning_tokens (21-524). See Business.not_good_enough.
the planted banking trap (13 of 36 applications, an otherwise-clean application with one unexplained mismatched bank_routing_number)
missed_banking_mismatch on the 13 trap applications: 76.9% (note-presence floor, 10 of 13 misread as explained) -> 0% (the fast tier) on the identical applications
measured
results/eval-b000-notepresence.json vs results/eval-r001-mp-onboard.json, same 36 applications, one variable (which decider reads them).
whether any submission_note is present at all, versus whether it specifically names the field
over_explained on the 25 unexplained-mismatch cells: 84.0% (note-presence floor, 21 of 25 misread) -> 0% (the fast tier) on the identical cells
measured
results/eval-b000-notepresence.json vs results/eval-r001-mp-onboard.json, same 25 cells.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
field accuracy
nothing yet
field answered
anything below 100%, on any run -- it means a reply was cut off or unparseable
missed banking mismatch
any nonzero value, on any run
over-explained
any nonzero value, on any run
per-field accuracy
any field falling below 100% on a re-run
reasoning-token share of output
nothing yet
latency
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
any call reaching MAX_TOKENS (3000) -- that is a truncated answer, not a long one
NextThe three you would add first
A live, code-level consistency check: recompute each field's match/mismatch call by comparing the declared value against its document value directly (string equality, the same comparison tools/build_corpus.py::derive_flags() already makes), and flag any disagreement with the model's own call before the UI renders it.This kit's own scorer (evals/scoring.py) already does this offline, against gold; nothing does it live, against the model's own reply.
A live check that the drafted summary names every field the model itself flagged mismatch or mismatch_explained, so a summary that undercounts its own flags is caught before a reader sees it.Nothing today cross-checks the free-text summary against the structured flags it is supposed to describe.
A red-team run against the submission_note field, the way fin-close attacked its own basis document.In a real deployment submission_note arrives from the applicant this kit's architecture treats as trusted with no verification step -- unmeasured for this kit (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the thinking setting in src/onboard.py.
What this cannot tell you
Whether the same 100%/0% figures hold on a second run, or on an application combining more than one of the five named patterns at once.
Whether a live, code-level consistency check (see add_first) would ever have caught a real disagreement -- this run's own model never produced one to test against.
Whether the prompt-only banking-trap and over-explaining rules hold against a submission_note crafted to mimic this kit's own explanation phrasing -- no red-team run exists for this kit.
Whether the same figures hold with reasoning explicitly disabled, matching the live app's own setting -- see Business.not_good_enough.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is an empty file with a comment explaining that the emptiness is load-bearing. The whole cross-check decision is three files: src/prompt.py, src/onboard.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
36 applications, generated from a fixed seed, never fetched. Deriving gold from the same declared-plus-document record (never the pattern name that seeded it) via derive_flags()/derive_outcome() -- with an assert inside build_application() catching any drift between intent and derivation before the file is even written -- is what keeps the internal-consistency check honest -- see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the seven field-pair taxonomy, the three-flag vocabulary and the output schema are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. Passing thinking through is a kwarg on one call site, not a client-library upgrade.
evaluation
evals/scoring.py
eval harnesses
exact match over a three-value flag vocabulary across seven fields, plus two named expensive-direction counts, is a nested loop and a couple of counters, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per application -- load, assemble, prompt, call, parse -- with no branching and no state carried between applications. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
An eighth declared field this kit's fixed seven-field FIELDS tuple does not name needs a hand-written FIELD_SOURCES/FIELD_MEANINGS entry plus new corpus template logic in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here.
No built-in output validation beyond src/prompt.py::parse()'s own tolerant-but-not-creative JSON parse -- a framework with a schema-and-consistency-check layer baked in might catch a flag that disagrees with the declared/document values, or a summary that disagrees with its own flags, before it renders; this kit's stdlib-only design did not build one (see Guardrails.fails).
What we could NOT verify
No port to any framework was actually built, so the comparison above is reasoning about the seams, not a measured alternative implementation.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-mp-onboard on the fast tier, 2026-08-20. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
3,139 ms
not yet known
nothing yet
Model, p95
5,039 ms
not yet known
nothing yet
Input tokens
39,006
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
12,336
not yet known
any call reaching MAX_TOKENS (3000) -- that is a truncated answer, not a long one
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-20, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
applications and their document blocks
data/applications.jsonl -- 36 applications, fixed seed, no clock read; your disk
one whole application record -- declared fields plus all three documents -- sent in the one call; src/app.py and src/onboard.py both read it by application_id, never modified after generation
gold flags and verification outcome
data/gold.jsonl -- 36 rows, computed by tools/build_corpus.py's derive_flags()/derive_outcome() by re-reading the same generated record, never carried over from the pattern name that seeded it
never -- scoring is in-process in evals/scoring.py, no judge model. src/onboard.py::load_gold()'s own docstring: NEVER read by check().
the key
.env at the repo root, shared by every kit -- never committed
only inside the Authorization header
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/check handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 36 applications and their gold flags -- byte-identically from a fixed seed (SEED=20260819) every time it is run, with an independent --verify pass re-deriving every gold row from the application records actually written to disk.
29,093 bytes, one file (data/applications.jsonl), 36 applications -- see Data.corpus. --verify confirmed 0 drift across all 36 rows this session. (lenses.Data.corpus, tools/build_corpus.py)
a real onboarding queue's own document formats and note phrasing are continuous, not a fixed seed, and a document too degraded to OCR cannot be recognised as such by this kit at all (see Data.breaks_on) -- neither was measured
point tools/build_corpus.py at your own businesses and document values, and every published field-accuracy/missed_banking_mismatch/over_explained figure is void -- they are this corpus's own phrasing templates and trap design (see Data.bring_your_own_boundary), not a property of the model
model
The seven field-pair taxonomy, the three-flag vocabulary and the output schema are declared once in src/prompt.py and sent in full inside the system prompt on every call -- a FIXED rulebook that does not vary by application. One call per APPLICATION carries the declared record and all three documents alongside it, behind src/adapters/__init__.py, reasoning left at the provider's default (on) -- the configuration r001-mp-onboard ships, and a DIFFERENT configuration from the one src/app.py's live UI actually runs (thinking explicitly off).
The rulebook portion of the system prompt is 887 (proportional-share) of 1,086 input tokens on the published call (lenses.LLM.prompt_parts[0]), sent identically on all 36 calls in r001-mp-onboard -- see LLM.prompt_verbatim for the exact text. 36 of 36 applications in r001-mp-onboard returned a reply that parsed cleanly (0 failures, finish_reason 'stop' on all 36), with the largest single call using 643 of the 3,000-token MAX_TOKENS budget (524 of it reasoning) -- real headroom. (lenses.Business.not_good_enough, results/eval-r001-mp-onboard.json, src/onboard.py)
a declared field this kit's fixed seven-field vocabulary doesn't name needs a hand-written FIELD_SOURCES/FIELD_MEANINGS entry plus new corpus template logic in tools/build_corpus.py -- never learned from data alone, and never measured here. Separately, a call needing meaningfully more reasoning than this run's worst case (524 tokens) was never measured against MAX_TOKENS=3000.
a different field set (added, removed or repaired declared fields) invalidates the whole eval at once -- gold and grading in evals/scoring.py are both keyed to this exact seven-field vocabulary. Verdicts are also per-model and this configuration ran once -- the live app's own thinking-off setting has never been scored against this corpus; see Business.not_good_enough.
labels
data/gold.jsonl, 36 rows -- the seven field flags and verification_outcome are all DERIVED from the same generated declared-plus-document record the model reads by derive_flags()/derive_outcome(), never from the pattern name that seeded the generator, and independently re-derived by --verify against the application records actually written to disk.
8 clean / 13 banking_trap / 8 explained_single / 4 escalated_nonbanking / 3 denied_multi gold applications over 36 total; 13 planted banking-trap applications, 25 unexplained-mismatch cells across all patterns -- see Eval.dataset and Eval.taxonomy. (lenses.Eval.dataset, lenses.Eval.taxonomy, mp-onboard-2026-08-19-36applications)
your own onboarding queue: hand-label the gold, which is the real work -- this kit's gold is a luxury of controlling the generator, and hand-labelled gold has an error rate this kit has never measured
field accuracy and missed_banking_mismatch/over_explained figures over this set reflect ONE planted trap family (a lone unexplained banking mismatch on an otherwise-clean application) and ONE assumption (every document already arrives text-extractable) -- a real onboarding queue's other failure shapes (see Data.breaks_on) are untested
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
an application where every field matches except bank_routing_number, unexplained, and the flag comes back anything other than mismatch
the banking trap this kit's whole design points at -- most often a reader that lets five or six clean fields talk it out of the one that isn't; the free note-presence floor misses 10 of 13
re-read the field on its own merits, independent of how clean the rest of the application looks, until Guardrails.add_first's live check exists (lenses.Eval.taxonomy (banking_trap) and example_row, r001-mp-onboard)
a generic or unrelated submission_note treated as covering a real discrepancy
the over-explaining failure -- a reader that treats 'a note exists' as 'the note explains this field'; the free floor over-explains 21 of 25 unexplained cells this way
check whether the note specifically names the field's discrepancy, not merely whether a note is present (lenses.Eval.baseline_note and data/SOURCES.md, b000-notepresence)
No machine symptom — this failure leaves no trace in any output.
a reply that ran out of tokens looks EXACTLY like a model that missed a mismatch. This run's largest call used 643 of 3,000 MAX_TOKENS, real headroom, but three sibling kits tonight found exactly this failure at a smaller ceiling. The control is recording finish_reason and reasoning_tokens on every call and reading them BEFORE believing a failure rate (lenses.Business.not_good_enough)
Concurrency and GPU sizing -- one serial call per application, nothing measured past 36. Provider-side retention -- provider-dependent, and on this kit the payload is application text naming businesses, owners and bank details, so that unknown IS the posture question. Where an application stops fitting the reasoning budget: this run's worst call spent 643 of 3,000 output tokens, so the ceiling was never approached, let alone found. The reasoning-off configuration, which matches the live app's own setting but has never been scored. Whether a document too degraded to read (a poor OCR pass) is handled sensibly -- every document in this corpus is clean text by construction. And what a long run costs: there is no cross-application deduplication and a run that dies at application 20 starts again -- no checkpointing, no resumption.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Field-pair exact match, plus banking-mismatch and over-explanation checks
Check a marketplace seller's application against their documents
PresenterOpens the private repo. Visible to admins only.
In one lineField-pair exact match, plus banking-mismatch and over-explanation checks
Does each of the seven field pairs' match/mismatch/mismatch_explained call match gold? Does the model avoid calling a genuinely mismatched bank_routing_number anything other than mismatch on the 13 applications where it is the ONLY field that disagrees (missed_banking_mismatch)? Does it avoid calling a genuinely unexplained mismatch mismatch_explained on any of the 25 cells across all patterns where gold is mismatch (over_explained)?
$0.00per 1,000 applications
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Application
APP-2024-1871
Business
Larkspur Pet Supply Co.
Pattern
banking_trap
Submission note (verbatim)
Let me know if you need any additional documents.
Declared routing number
137760256
Bank letter routing number
451516861
Gold flags
all seven match except bank_routing_number=mismatch
Note-presence floor's bank_routing_number call
mismatch_explained (WRONG -- the note is present but names nothing about banking)
The model's bank_routing_number call
mismatch
The model's drafted summary
The only flagged inconsistency is bank_routing_number: the application declares 137760256 while the bank verification letter lists 451516861. The submission note does not address this discrepancy.
Scored as
correct -- the trap the free note-presence floor fails (a generic note wrongly read as an explanation) and the fast tier does not
Grader
Verdict
Why
Field-pair exact match, plus banking-mismatch and over-explanation checks
correct
APP-2024-1871: all seven fields match their supporting documents except bank_routing_number -- the application declares 137760256 while the bank verification letter lists 451516861. The submission note reads 'Let me know if you need any additional documents.': present, but naming nothing about banking. Gold: bank_routing_number mismatch, all six other fields match, verification_outcome denied. The free note-presence floor sees a note attached and calls the mismatch explained -- a false-clean verdict on a fraudulent-pattern application. The fast tier read the note against the field, correctly called it a plain mismatch, and its drafted summary named exactly this discrepancy and nothing else.
The formulaWhat it computes
field_accuracy = correct field flags / answered, over answered=252 (36 applications x 7 fields). missed_banking_mismatch = an application where gold's bank_routing_number is mismatch, all six other fields are gold match, and the model's own bank_routing_number flag is anything other than mismatch, over the 13 applications matching that shape. over_explained = a cell where gold is mismatch and the model's flag is mismatch_explained, over the 25 cells across all patterns where gold is mismatch.
The analysisWhat it actually did
Model
Result
the fast tier
100.0% field accuracy
In operationWhat to monitor
Reference standard: tools/build_corpus.py's derive_flags()/derive_outcome(), which compare each declared value against its document value and scan submission_note for field-specific trigger phrases -- re-derived independently via --verify against the applications actually written to disk, asserted zero drift.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own trap design and note templates, which come from tools/build_corpus.py's fixed vocabulary.
Watch these
missed_banking_mismatch as a raw count, never folded into field_accuracy -- it is the expensive-direction error this kit's trap exists to catch
over_explained specifically on applications with ANY submission_note, since a reader that treats note-presence as explanation fails exactly there (see baseline_note)
answered vs asked -- 100% this run, but a truncated reply answers fewer fields while every field it DID answer stays correct, the same truncation-artefact shape sibling kits found at a smaller MAX_TOKENS
Alarm on
Any nonzero missed_banking_mismatch or over_explained, on any run. Both zero this run.
How tight can the band be? There is no tolerance band -- every field is an exact match over a three-value vocabulary. Nothing here is a continuous quantity to round.
Cadence: Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS.
The decisionWhen to reach for it
Use it
Gold is DERIVED from the same generated declared-plus-document record the model reads -- true of every kit corpus, never true of a real verification team's own applications.
Do not use it
The truth is not known in advance -- the normal state of a real onboarding queue, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning