The business caseThe problem this solves
A delivery marketplace settles a restaurant's orders and sends a statement: the menu subtotal it took, the promotions it booked against the store, the delivery and service fees it charged, and the tax it says it collected and settled. Whether the marketplace or the store is the party that must remit that tax is not on the statement — it is a fact about the jurisdiction, and on 51 of these 63 packets the platform's own taxable base is built on something that is not what the columns say it is: an order cancelled and refunded, one basket settled twice, a promotion the marketplace funded and rebilled to the store, an exemption on a certificate that has expired. Somebody has to open the packet, read the note under every line, decide which orders are real, notice whether a rate change actually took effect, and say where the period really stands before anybody files anything. Opening one settlement packet, adding up the menu subtotals and the two fee columns by hand under whichever fee flags the supplied rule row carries, reading each note under an order line to decide whether the order really happened and who funded its promotion, checking the operator notes for a rate change that actually took effect, taxing the base, comparing it with what the platform says it settled, and deciding whether the difference breaks either investigate threshold.
Audience
A franchise group's indirect-tax desk working a monthly marketplace settlement, and the finance controller behind it. Whoever decides that this period is filed as the platform reports it, goes to the marketplace as a query, or waits for somebody to supply a rule row. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual settlement packet
The corpus is 63 settlement packet, 0.19 MB (txt 63). It is generated because it has to be. A real marketplace settlement statement is a franchisee's own trading record, and the exact shapes this kit measures — an order cancelled after settlement, a promotion rebilled to the store in error, an exemption certificate that names somebody else — are the lines a restaurant group would least want published. Generating it also makes the key DERIVED rather than written: each packet is built as a structure, the panels are rendered from it, and MTR-2026 is applied to the same structure by src/policy.py. There is no second place the answer lives. ⚑ AND IT IS WHY THE RULE TABLE COULD BE AN INPUT AT ALL: the jurisdictions are invented, so nothing here needed a real facilitator status to be researched, asserted or believed — which is precisely the block this use case was recorded as being under.
The corpus
- The 63 settlement packetgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every store and marketplace is an invented trading name, every store number, statement number, jurisdiction code, order reference and certificate id is arithmetic on the packet index, and there is no personal data at all — evals/check_labels.py sweeps all 63 packets for five families of identifier on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your settlement packet. That is the whole change — there is no database to migrate.
================================================================================================
MARKETPLACE TAX SETTLEMENT PACKET MTX-0001
Store: STR-4100 - Northgate Crossing (invented)
Statement: STM-410007 Marketplace: Sablewood Eats (invented) Period: 2026-07-01 to 2026-07-31 Procedure: MTR-2026
================================================================================================
STORE AND PERIOD AS THE FILING MASTER HOLDS IT
jurisdiction JD-4200
filing entity ENT-2200
tolerance pct 0.25 pct of tax due
threshold tax 75.00
threshold pct 4.00 pct of tax due
period 2026-07-01 to 2026-07-31
FACILITATOR RULE TABLE AS SUPPLIED
JD-4200 6.000 facilitator remits: no delivery fee: exempt service fee: taxable
JD-4213 6.750 facilitator remits: yes delivery fee: exempt service fee: exempt
JD-4226 7.125 facilitator remits: yes delivery fee: taxable service fee: exempt
JD-4300 6.375 facilitator remits: no delivery fee: taxable service fee: taxable
ORDER LINES AS THE PLATFORM SETTLED THEM
ORDER DATE SUBTOTAL DISCOUNT DELIVERY SERVICE TAX KIND STATUS REMITTER REF MEMO
ORD-0001 2026-07-02 18.00 0.00 3.99 2.50 1.23 order SETTLED PLATFORM REF-10000 standard delivery order
ORD-0002 2026-07-04 29.00 6.00 4.99 3.25 1.58 order SETTLED PLATFORM REF-10001 promotion booked to the store, 3 items at 2.00 each
TIP-0090 2026-07-04 3.00 0.00 0.00 0.00 0.00 tip INFO - REF-70000 gratuity passed through to the courierAbridged — the file continues.
The outcomeWhat a good result looks like
One packet in, one row out: which order lines this procedure treats differently from the statement that printed them, whether a combined-rate change was in force, and one verdict from a closed set of six — NO-RULE, BASE-BROKEN, ESCALATE, UNDER-REMITTED, OVER-REMITTED or RECONCILED. The taxable base, the tax due, the tax settled, the difference and WHICH SIDE OWES IT are re-derived in pure code from those readings and the supplied rule row, so they follow MTR-2026 whatever the reply said.
And when it cannot
And what it does when it cannot. On the scored run 63 of 63 replies parsed and nothing stopped at the ceiling, so there is no unparsed row to report. What it gets WRONG is published by name. ⚠︎ THE MOST IMPORTANT ONE IS NOT AN ACCURACY FIGURE: on all 3 packets whose jurisdiction the supplied rule table does not carry, the arm answered a verdict anyway — UNDER-REMITTED, OVER-REMITTED and RECONCILED — although the prompt says in terms that the answer is NO-RULE. The prompt did not hold that cap; R-6 in src/policy.py did, and the published column is NO-RULE on all three. It is counted as classified_without_a_rule rather than folded into a score.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You have a verified, jurisdiction-by-jurisdiction facilitator table already — this kit, and point the register at your table
The table is an INPUT here, not a belief. Every committed run re-scores against a new one for $0.00, and a jurisdiction you have not verified is simply absent. - Your statement's bad lines announce themselves — every cancellation says "cancelled", every funded promotion says "marketplace" — the free rules floor, and do not buy a call at all
A keyword list over the notes reaches everything a keyword can reach, for $0.00, and it gets 30 of 63 packets whole. - Your operator notes routinely describe a cancellation, a refund or a funded promotion OF SOMETHING ELSE — the paid call, and read the keep-note family's own rate first
The floor gets that family 0 of 16 because a keyword rule cannot tell which line a sentence is about. The paid call gets 11. - You need the taxable base and nothing else — either arm — and derive the figure in code from whichever reading you have
The arm's OWN sum is right on 4 of 63 and the free floor's on 34. Put either arm's cited set through thirty lines of integer arithmetic and the paid one gives 50. The arithmetic is not what a model is for.
And where nothing here is good enough:
- Somebody downstream files a return on the output — neither, until that changes
This kit never files, never remits, never claims and never disputes, and it never states a jurisdiction's tax position as a fact of law. It reconciles a period against a rule table somebody else verified and stops.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own settlement packets in the same shape and data/jurisdictions.json with your own filing master AND YOUR OWN VERIFIED FACILITATOR RULE TABLE, then rebuild the key by labelling them. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Asking any model — this one included — what a jurisdiction's facilitator status is. On the 3 packets where the table was silent, this arm answered a verdict every time. That is the case against the best-fitting scenario (“You have a verified, jurisdiction-by-jurisdiction facilitator table already”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A statement that is not fixed-width columns. src/statement.py's line regex is the shape these packets print; a CSV or a marketplace API export needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-marketplace-tax. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every screenshot. python3 -m evals.check_labels and python3 tools/build_corpus.py --check both run on a machine with nothing installed.






