The business caseThe problem this solves
When a client moves an account to a new firm, the positions arrive first and the COST BASIS arrives separately, later, and sometimes not at all. Until it lands — and lands correctly — the receiving firm cannot tell a client what a sale will cost them, and at tax-form production the missing figure stops being an operations problem and becomes a client-facing one. Meanwhile the gaps get filled by hand: somebody keys a number in from a client's old statement or a phone call, and that manual override is now a basis the client will be taxed on with nothing behind it but a name in a field. The team's own words for the job are three questions — show me every lot with missing or implausible basis now rather than after client complaints; show me every manual override, its evidence and the named supervisor who approved it; show me the delivering firms with the worst basis quality and how old the chase list against them is. Working a transfer basis queue lot by lot out of two screens — what the delivering firm sent and what the receiving book carries — subtracting the two by hand, eyeballing whether a per-unit figure is sane, and remembering which overrides still need a supervisor and which delivering firms owe an answer to a chase raised weeks ago.
Audience
A cost basis / account transfers analyst working a transfer queue, and the tax operations supervisor who owns the overrides they raise. The split between them is the point of the kit and is never relaxed: an analyst may chase a delivering firm and work a routine lot; only a named supervisor approves a manual override, at any size. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual transferred-in lot review packets
The corpus is 64 transferred-in lot review packets, 0.08 MB (txt 64). It is generated because it has to be. A real transferred-in basis queue is client tax data: account numbers, holdings, acquisition dates and the cost figures a person will be taxed on. There is no public corpus of it and there should not be. So the generator builds 64 packets across 21 case families from one seed, renders each one, and DERIVES the answer key from the same structure it rendered from — the key is never typed. tools/build_corpus.py --check rebuilds packets, key and register and diffs them, and it reproduces byte-identically under different PYTHONHASHSEEDs.
The corpus
- The 64 transferred-in lot review packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md carries the seed, the licence pointer, the case-family map and the generator's own attack on itself.
Swap this folder for your own material and the kit is pointed at your transferred-in lot review packets. That is the whole change — there is no database to migrate.
COST BASIS REVIEW PACKET BGL-0001
Prepared 2026-09-02 under BGR-2026 | Transferred-in lot basis review
LOT FACTS
Receiving account ACCT-48733
Security Wenlock International Value Fund
Symbol WKIVX
Quantity 1,275.0000 units
Delivering firm Ardenvale Brokerage
Internal tax-form cutoff 2026-10-31
Packet reference BGL-0001
DELIVERED BASIS RECORD
Record status PROVIDED
Total cost transmitted $ 343,496.80
Acquisition date transmitted 2016-03-04
RECEIVING BOOK
Total cost carried $ 343,497.75
Acquisition date carried 2016-03-04
Reference price on that date $ 104.83 per unit
OVERRIDE RECORD
none entered
CHASE LOG
no request has been sent to the delivering firm for this lot
NOTES
Transfer operations confirm the position itself settled cleanly; only the basis record is in question.
The relationship manager has seen the transfer summary and had nothing to add for this lot.
The outcomeWhat a good result looks like
One lot in, six graded answers out: the gap to the cent (null where nothing was transmitted, which is 22 of the 64 lots and a real answer rather than a blank), the basis finding, what the override evidence actually is, the lot's position under BGR-2026, which list it goes on today, and the one line in the packet the reading came from. On the scored run that was 378 of 384 graded cells as answered and 384 of 384 once the rulebook is re-applied in code — against 357 of 384 for a free phrase-and-arithmetic floor and 175 of 384 for a floor that reads nothing.
And when it cannot
And what it does when it cannot. The paid call dropped the referral on 3 of the 3 lots whose override was approved by somebody who is NOT on the supervisor roster — it read the evidence correctly every time and then answered NO-ACTION, which is the single most consequential mistake available on this corpus: an unapproved override stays on the book and the supervisor never sees it. All 3 come back once src/policy.py checks the approver against data/supervisors.json, because that is a FILE and not a sentence. ⚠︎ READ THAT AS THE KIT'S CENTRAL RESULT, NOT A FOOTNOTE: the model alone is not the product here, and a deployment of the reading without the pure-code station would have shipped those 3.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You want the gap to the cent and the basis finding, and nothing else — the free rules floor alone — evals/baseline.py
64 of 64 gaps exact and 64 of 64 findings for $0.00, tying the paid call on both. A real transmitted-basis record is fixed-layout too. - You want to know what is actually behind your manual overrides — the paid call, and read the override_evidence field on its own
64 of 64 against the floor's 53. That single field is the entire measured margin: on the 39 lots with no override the two arms are identical. - You want the queue decision itself to be defensible — the pack as shipped — the call plus src/recheck.py
64 of 64 lots against the raw call's 61. The three it rescues are all the same class: an override approved by somebody not on the roster. - Your book has few manual overrides — the free floor, and re-measure before buying anything
Every cell the call buys on this corpus is inside the 25 override lots. With fewer of them there is proportionally less to buy, and this kit's margin does not transfer.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own rendered packets, data/lots.json with your own transfer register rows (one per packet), data/supervisors.json with your own roster and data/policy.json with your own rulebook and thresholds. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for arithmetic. It is most of the work in this job and none of the value. That is the case against the best-fitting scenario (“You want the gap to the cent and the basis finding, and nothing else”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A scanned or photographed statement. Every arm here — the model's reading, the floor's arithmetic and the citation locator — is bound to a rendered text layout, and nothing in this kit does OCR. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the labelled line is the line an analyst would have quoted. The key names ONE line per reading and a reviewer might accept a neighbouring one carrying the same fact; this key would score that zero. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-02 — r001-basis-gap. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9269 and scores every graded cell on both free floors, the committed scored run and the committed adversarial run, for $0.00 and with no network. Nothing needs installing: standard library end to end.



