The business caseThe problem this solves
A franchise finance desk receives one period's sales report from a franchisee and has to decide whether the royalty remitted matches what the agreement's own ladder produces from the point-of-sale columns behind it. The columns are printed on the packet. The trap is not the arithmetic: it is whether each column may be USED at all for this period, which is settled six lines deep in a register control trail written in prose — a terminal journal withdrawn and restated, a delivery channel feed confirmed for the wrong period, a discount schedule whose funding side was never returned. Reading one packet's six dated register control notes against the agreement's reporting terms to decide which point-of-sale columns are usable for this period, then running R-2 to R-10's exclusion ladder in cents and comparing the result with what was remitted.
Audience
A franchise finance analyst reconciling one filed period, and the franchise business manager who decides whether to open a variance conversation with the operator. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual period sales reconciliation packets
The corpus is 64 period sales reconciliation packets, 0.29 MB (json 3 · jsonl 1 · md 1 · txt 64). Because the shape of a reported-sales failure is not the arithmetic, it is whether a printed column may be used at all. Every packet carries all six point-of-sale columns and a six-note control trail, so a reader that never opens the trail still has a complete-looking set of numbers — which is exactly how a wrong reconciliation gets produced confidently. ⚑ AND THE NATURAL ERROR CANCELS ON THE FAMILY THAT COMMITS IT: an arm that nets the delivery commission out reproduces the franchisee's own figure on the commission-netted packets, so the variance comes out zero and the period reads clean, while the same habit raises a false flag everywhere else. That is why the FIGURES are the headline and the status is not: a constant reply takes 42.19 pct of status having read nothing and 0.0 pct of the full row.
The corpus
- The 64 period sales reconciliation packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere. All 64 packets, data/period_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your period sales reconciliation packets. That is the whole change — there is no database to migrate.
BRACKENMOOR KITCHENS - FRANCHISE PERIOD SALES RECONCILIATION PACKET
SYNTHETIC SAMPLE DATA. GENERATED, NOT REAL. NO REAL FRANCHISOR, FRANCHISEE OR STORE.
Packet RSR-0001 Store STR-0717 Period PER-2026-W20 Submitted 2026-05-18
STORE AND AGREEMENT
Store reference STR-0717
Other store in the same trade area STR-0666
Franchise agreement BK-FA-2026
Royalty rate on net sales 5.50 %
Period start 2026-05-11
Period end 2026-05-17
Current terminal journal TJL-2026-5138
Secondary terminal journal TJL-2026-4484
Prior-period terminal journal TJL-2025-4292
Current tax table TXT-2026-33
Prior tax table TXT-2025-45
Current delivery channel feed DCH-2026-3336
Prior delivery channel feed DCH-2025-3775
Current promotion schedule PRM-2026-44
Prior promotion schedule PRM-2025-43
Current gift-card ledger GCL-2026-5213
Prior gift-card ledger GCL-2025-5665
REPORTED BY THE FRANCHISEE
Net sales reported $50,253.00
Royalty remitted $2,763.91
AMOUNTS FROM THE POINT-OF-SALE SYSTEM
POS gross rung, all channels $58,700.00
Sales tax collected $4,960.00
Third-party delivery gross, before commission $7,590.00
Third-party delivery commission retained $1,550.00Abridged — the file continues.
The outcomeWhat a good result looks like
Either a reconciliation carrying ten figures — net sales, royalty due, royalty remitted, variance, tolerance and the five exclusions the ladder takes out — each naming the BK-FA-2026 term that produced it, plus a route, a timing verdict and a repeat-variance verdict; or a refusal naming exactly ONE unusable point-of-sale column and carrying no figures at all.
And when it cannot
⛔ THE HEADLINE IS A LOSS AND IT IS PUBLISHED AS ONE. The measure is the FULL ROW — the status, the one named unusable column, all ten figures, all ten BK-FA-2026 term citations, the route, the timing and the repeat verdict, on one packet. Rechecked through src/recheck.py the paid call takes 45 of 64 (70.31 pct). The free floor of record — the printed point-of-sale columns read by pure Python, $0.00, no network — takes 53 of 64 (82.81 pct). McNemar exact on the same 64 paired items: only the paid call 6, only the floor 14, p = 0.115318 — NOT SIGNIFICANT, and the point estimate runs AGAINST the call. The honest sentence is that this corpus cannot separate the paid call from its free floor on the full row.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- every point-of-sale column on the packet is usable and printed — the free printed-columns floor, evals/baseline.py --floor fields
it takes 53 of 64 full rows at $0.00 and the paid call takes 45. There is nothing here for a model to add and the measurement says so. - the agreement's rules can be retyped in code and kept in step — the free ceiling arm, evals/baseline.py --floor full
64 of 64 at $0.00. It beats the paid call by 19 packets at p = 0.000004. - the answer turns on a control trail written in prose that nobody will retype — the paid call, with src/recheck.py in front of every figure it returns
on the 11 reading-required packets it takes 6 where the floor of record takes 0. ⚠︎ AND A FREE WORD LIST TAKES 5 OF THE SAME 11 — the margin is one packet.
And where nothing here is good enough:
- you want the arm's own arithmetic on the page — nothing here
as answered the call takes 5 of 64 full rows and gets its own ladder wrong on 47 of 64. Every figure this kit publishes is re-derived in code.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered period packets and data/period_records.json at their structured twin, keep the six column names and the note shape, and every arm runs unchanged: python3 -m evals.run --floor fields scores the free floor of record on your packets at $0.00 before you buy a single call. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call whose reading you do not need. That is the case against the best-fitting scenario (“every point-of-sale column on the packet is usable and printed”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet whose register control trail contradicts itself — two notes of equal date, one withdrawing a journal and one restating it. The corpus carries the near miss (a withdrawal and a live restatement of DIFFERENT figures on the same journal id, RSR-0054 and RSR-0064) and nothing resolves a true tie. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether a second run of the same 64 packets would move the full row at all — one scored run was bought and no repeat was. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-reported-sales. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on a clean checkout with no key configured and no network: python3 -m evals.run --floor fields scores all 64 packets in 0.5 s at $0.00, and python3 -m src.app --port 9494 serves the whole board — every committed arm, both columns — with the reconcile button disabled and the reason printed beside it. Two minutes from clone to a scored free floor, and nothing is bought to get there.








