The business caseThe problem this solves
An append-only record answers two different questions and most systems only answer one of them. “What is the billed total on this order?” has an answer today; “what was it on 8 March?” has a different one, because entries recorded after 8 March are not part of what the record held on 8 March, and entries recorded before it that only became effective later are not part of it either. Today somebody opens the log, reads the two dates on every line by eye, decides which lines the record actually carried on the day in question, and then has to say in writing what about that answer a reader should not trust — a correction that landed late, an amendment already in flight, a stream they were not allowed to read on that date. The arithmetic is easy and nobody gets it wrong. The writing is where the answer is lost. Reading an append-only log line by line against a date — deciding which entries the record actually held then, which the asker was allowed to see then, and then writing the qualifications a reader needs — and replaces only the LAST of those three. The fold, the scope test and the caveat derivation are pure code and cost nothing; the call is bought for the writing.
Audience
Whoever has to answer a point-in-time question in writing and sign it — a records desk, a dispute analyst, an auditor's correspondent. The decision they are making is not what the figure is: free code already produced that, exactly, and this kit never asks a model for it. The decision is WHICH of the record's own caveats a reader needs in order to trust the figure, and in what order. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual as-of question packets
The corpus is 64 as-of question packets, 0.15 MB (json 4 · jsonl 1 · md 2 · txt 64). Because a point-in-time answer fails on WHICH QUALIFICATIONS ARE SAID, not on the arithmetic — so the corpus had to make the arithmetic free and the choosing hard. Every packet carries an append-only log whose entries each have a recorded date and an effective date, and 36 of the 64 carry a TWIN PAIR: a moving entry and a non-moving one that share their head clause word for word, sit on the same subject, the same line, the same amount and the same pair of dates, and differ only in a trailing clause. Five neutral tails leave an entry moving and five nullifying tails do not, and all ten verbs — billed, corrected, raised, voided, marked — appear on both halves, so no word list can separate them. 224 caveats are derived across the corpus and 84 of them bear on the question actually asked; 25 packets cannot fit all of theirs inside the answer budget. That last number is the kit.
The corpus
- The 64 as-of question packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your as-of question packets. That is the whole change — there is no database to migrate.
AS-OF QUESTION PACKET aq-0001
==============================================================================
QUESTION
Subject ORD-4400
Account AC-100
Asked What was the billed total on ORD-4400?
As of 2026-03-08
Asked by desk 03
Record taken 2026-09-01
READING GRANTS ON FILE
GR-0201 desk 03 BILLING from 2026-01-13 open-ended
GR-0202 desk 03 STATUS from 2026-01-13 open-ended
GR-0203 desk 03 CREDITS from 2026-01-13 open-ended
GR-0204 desk 03 BILLING from 2025-12-04 to 2026-01-12
EVENT RECORD, APPEND-ONLY
id recorded effective stream
EV-0001 2026-02-02 2026-02-02 STATUS
ORD-4400 is marked OPEN, and the entry is reissued the same day.
EV-0002 2026-02-03 2026-02-03 BILLING
Line L1 of ORD-4400 is billed at $2,105.31 on invoice IN-9902, and the entry is reissued the same day.
EV-0003 2026-02-05 2026-02-05 BILLING
Line L2 of ORD-4400 is billed at $203.81 on invoice IN-9903, which the raising desk has acknowledged.
EV-0004 2026-02-07 2026-02-07 BILLING
Line L3 of ORD-4400 is billed at $297.34 on invoice IN-9904, and the entry is signed off that week.
EV-0005 2026-02-14 2026-02-14 CREDITS
Credit memo CM-0045 for $52.73 is raised against line L1 of ORD-4400, and the entry is signed off that week.
EV-0006 2026-03-05 2026-03-05 STATUS
ORD-4400 is marked SETTLED, with the supporting paperwork attached.
EV-0007 2026-03-11 2026-03-02 BILLING
Line L1 of ORD-4400 is billed at $1,388.84 on invoice IN-9907, and the entry is reissued the same day.
EV-0008 2026-03-11 2026-03-02 BILLINGAbridged — the file continues.
The outcomeWhat a good result looks like
Free code folds the record as of the date asked about and derives every caveat it raises. The one call receives those caveats as sentences with their record ids and nothing else — no kinds, no streams, no line numbers, no materiality flags and no figure — and writes the answer. Two set differences grade it: invented = said − derived is every fact asserted that the fold did not derive, and missed = material − said is every caveat the fold marked as bearing on this question that the answer left out. A packet is EXACT when both are empty. On the 25 CROWDED packets — the ones where restating every derived caveat overruns the 560-character answer budget, so the arm has to choose — the paid call is 23 of 25 against the free bar's 17 and the floor of record's 12.
And when it cannot
⚠︎ THE HEADLINE IS A PAIR AND IT IS SPLIT; IT IS NEVER AVERAGED. The paid call wins the choosing cell 23/25 against the bar's 17/25 (exact McNemar on the registered slice: 6 packets to 0, p = 0.031250) AND LOSES THE GUARDRAIL CELL: invented_clean is 56 of 57 on both paid runs against 57 of 57 on every free arm. Nothing that only restates handed sentences can invent a fact, so the free arms hold the best possible value there by construction, and the paid call breaks the tie downward. It also writes past the budget on 11 of 57 packets where no free arm ever does, and the truncation is the whole of both error columns.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A desk whose point-in-time answers are short and rarely carry more than a couple of qualifications — the free bar, and do not buy a call
the floor of record is already 32 of 32 on the roomy packets for $0.00, and the bar adds five more crowded ones for the same money - A desk where the record routinely raises more qualifications than an answer can hold, so somebody has to choose — this kit, with the paid call, beside the free bar on every packet
that is the only cell the money moves — 23 of 25 against the bar's 17, 6 discordant packets to 0 at p = 0.031250 - A desk that must be able to prove nothing was invented — the free bar
every free arm is 57 of 57 oninvented_cleanby construction — nothing that only restates handed sentences can invent a fact — and the paid call is 56 - A desk whose askers have time-boxed reading grants — this kit, for the scope half rather than the writing half
the grant windows are tested against the date asked about and never against today, the refusal is the product, and the withheld-stream caveat names the stream and no figure — all in code, all free
And where nothing here is good enough:
- A desk that wants a number for whether the WORDING is good — neither arm — this kit does not measure that
two set differences measure which caveats were said, not whether the sentence reads well or puts the right one first
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered packets and write the reader in tools/build_corpus.py for your log's shape, then retype your own rulebook in src/rules.py: the streams, what each question depends on, the refusal grounds, the caveat kinds and the materiality test. The boundary is the ANSWER KEY, not the packets. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a cell that has no headroom — the roomy half of this corpus buys literally nothing. That is the case against the best-fitting scenario (“A desk whose point-in-time answers are short and rarely carry more than a couple of qualifications”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A log whose entries carry only ONE date. The entire kit is the difference between recorded and effective; with one date there is no fold, no late-recorded caveat and no not-yet-effective caveat, and the floor of record answers everything for $0.00. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether any of this holds on a real organisation's records. Every figure is against a key derived from this kit's own generator at seed 20260921, on an invented rulebook, with a retail order-and-billing story as the worked example only. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-21 — r001-asof-record. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on this repository with no key configured: python3 -m evals.baseline printed all eight free arms over 64 packets in 0.09 seconds at $0.00, and the board came up on the no-key port with the one billable button really disabled. Because this kit SHIPS its reply caches, the same clean checkout also re-scored both paid runs to the digit — 55/57 and 53/57 exact — with no key and no network. What it could not do is buy a new answer.














