Home › Use Cases › Answer one question as of a past date, and say what the record had not yet recorded
Use caseUC0523
🧪 Use-case kit · runnable

Answer one question as of a past date, and say what the record had not yet recorded

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An append-only record answers two different questions and most systems only answer one of them. “What is the billed total on this order?” has an answer today; “what was it on 8 March?” has a different one, because entries recorded after 8 March are not part of what the record held on 8 March, and entries recorded before it that only became effective later are not part of it either. Today somebody opens the log, reads the two dates on every line by eye, decides which lines the record actually carried on the day in question, and then has to say in writing what about that answer a reader should not trust — a correction that landed late, an amendment already in flight, a stream they were not allowed to read on that date. The arithmetic is easy and nobody gets it wrong. The writing is where the answer is lost. Reading an append-only log line by line against a date — deciding which entries the record actually held then, which the asker was allowed to see then, and then writing the qualifications a reader needs — and replaces only the LAST of those three. The fold, the scope test and the caveat derivation are pure code and cost nothing; the call is bought for the writing.

Audience

Whoever has to answer a point-in-time question in writing and sign it — a records desk, a dispute analyst, an auditor's correspondent. The decision they are making is not what the figure is: free code already produced that, exactly, and this kit never asks a model for it. The decision is WHICH of the record's own caveats a reader needs in order to trust the figure, and in what order. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual as-of question packets

The corpus is 64 as-of question packets, 0.15 MB (json 4 · jsonl 1 · md 2 · txt 64). Because a point-in-time answer fails on WHICH QUALIFICATIONS ARE SAID, not on the arithmetic — so the corpus had to make the arithmetic free and the choosing hard. Every packet carries an append-only log whose entries each have a recorded date and an effective date, and 36 of the 64 carry a TWIN PAIR: a moving entry and a non-moving one that share their head clause word for word, sit on the same subject, the same line, the same amount and the same pair of dates, and differ only in a trailing clause. Five neutral tails leave an entry moving and five nullifying tails do not, and all ten verbs — billed, corrected, raised, voided, marked — appear on both halves, so no word list can separate them. 224 caveats are derived across the corpus and 84 of them bear on the question actually asked; 25 packets cannot fit all of theirs inside the answer budget. That last number is the kit.

The corpus

  • The 64 as-of question packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your as-of question packets. That is the whole change — there is no database to migrate.

One as-of question packet, as the model receives itaq-0001.txt · 1 of 64
AS-OF QUESTION PACKET  aq-0001
==============================================================================

QUESTION
  Subject            ORD-4400
  Account            AC-100
  Asked              What was the billed total on ORD-4400?
  As of              2026-03-08
  Asked by           desk 03
  Record taken       2026-09-01

READING GRANTS ON FILE
  GR-0201   desk 03   BILLING    from 2026-01-13 open-ended
  GR-0202   desk 03   STATUS     from 2026-01-13 open-ended
  GR-0203   desk 03   CREDITS    from 2026-01-13 open-ended
  GR-0204   desk 03   BILLING    from 2025-12-04 to 2026-01-12

EVENT RECORD, APPEND-ONLY
  id         recorded     effective    stream
  EV-0001    2026-02-02   2026-02-02   STATUS
      ORD-4400 is marked OPEN, and the entry is reissued the same day.
  EV-0002    2026-02-03   2026-02-03   BILLING
      Line L1 of ORD-4400 is billed at $2,105.31 on invoice IN-9902, and the entry is reissued the same day.
  EV-0003    2026-02-05   2026-02-05   BILLING
      Line L2 of ORD-4400 is billed at $203.81 on invoice IN-9903, which the raising desk has acknowledged.
  EV-0004    2026-02-07   2026-02-07   BILLING
      Line L3 of ORD-4400 is billed at $297.34 on invoice IN-9904, and the entry is signed off that week.
  EV-0005    2026-02-14   2026-02-14   CREDITS
      Credit memo CM-0045 for $52.73 is raised against line L1 of ORD-4400, and the entry is signed off that week.
  EV-0006    2026-03-05   2026-03-05   STATUS
      ORD-4400 is marked SETTLED, with the supporting paperwork attached.
  EV-0007    2026-03-11   2026-03-02   BILLING
      Line L1 of ORD-4400 is billed at $1,388.84 on invoice IN-9907, and the entry is reissued the same day.
  EV-0008    2026-03-11   2026-03-02   BILLING

Abridged — the file continues.

The outcomeWhat a good result looks like

Free code folds the record as of the date asked about and derives every caveat it raises. The one call receives those caveats as sentences with their record ids and nothing else — no kinds, no streams, no line numbers, no materiality flags and no figure — and writes the answer. Two set differences grade it: invented = said − derived is every fact asserted that the fold did not derive, and missed = material − said is every caveat the fold marked as bearing on this question that the answer left out. A packet is EXACT when both are empty. On the 25 CROWDED packets — the ones where restating every derived caveat overruns the 560-character answer budget, so the arm has to choose — the paid call is 23 of 25 against the free bar's 17 and the floor of record's 12.

And when it cannot

⚠︎ THE HEADLINE IS A PAIR AND IT IS SPLIT; IT IS NEVER AVERAGED. The paid call wins the choosing cell 23/25 against the bar's 17/25 (exact McNemar on the registered slice: 6 packets to 0, p = 0.031250) AND LOSES THE GUARDRAIL CELL: invented_clean is 56 of 57 on both paid runs against 57 of 57 on every free arm. Nothing that only restates handed sentences can invent a fact, so the free arms hold the best possible value there by construction, and the paid call breaks the tie downward. It also writes past the budget on 11 of 57 packets where no free arm ever does, and the truncation is the whole of both error columns.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A desk whose point-in-time answers are short and rarely carry more than a couple of qualifications — the free bar, and do not buy a call
    the floor of record is already 32 of 32 on the roomy packets for $0.00, and the bar adds five more crowded ones for the same money
  • A desk where the record routinely raises more qualifications than an answer can hold, so somebody has to choose — this kit, with the paid call, beside the free bar on every packet
    that is the only cell the money moves — 23 of 25 against the bar's 17, 6 discordant packets to 0 at p = 0.031250
  • A desk that must be able to prove nothing was invented — the free bar
    every free arm is 57 of 57 on invented_clean by construction — nothing that only restates handed sentences can invent a fact — and the paid call is 56
  • A desk whose askers have time-boxed reading grants — this kit, for the scope half rather than the writing half
    the grant windows are tested against the date asked about and never against today, the refusal is the product, and the withheld-stream caveat names the stream and no figure — all in code, all free

And where nothing here is good enough:

  • A desk that wants a number for whether the WORDING is good — neither arm — this kit does not measure that
    two set differences measure which caveats were said, not whether the sentence reads well or puts the right one first

At a glanceHow the whole thing runs

92%crowded exact pct
1,416 msp50, end to end
$11.87per 1,000 as-of question packets · Claude Fable 5

Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own rendered packets and write the reader in tools/build_corpus.py for your log's shape, then retype your own rulebook in src/rules.py: the streams, what each question depends on, the refusal grounds, the caveat kinds and the materiality test. The boundary is the ANSWER KEY, not the packets. Corpus lens →
When is this the wrong choice?Avoid: Paying for a cell that has no headroom — the roomy half of this corpus buys literally nothing. That is the case against the best-fitting scenario (“A desk whose point-in-time answers are short and rarely carry more than a couple of qualifications”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A log whose entries carry only ONE date. The entire kit is the difference between recorded and effective; with one date there is no fold, no late-recorded caveat and no not-yet-effective caveat, and the floor of record answers everything for $0.00. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real organisation's records. Every figure is against a key derived from this kit's own generator at seed 20260921, on an invented rulebook, with a retail order-and-billing story as the worked example only. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-21 — r001-asof-record. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Observed on this repository with no key configured: python3 -m evals.baseline printed all eight free arms over 64 packets in 0.09 seconds at $0.00, and the board came up on the no-key port with the one billable button really disabled. Because this kit SHIPS its reply caches, the same clean checkout also re-scored both paid runs to the digit — 55/57 and 53/57 exact — with no key and no network. What it could not do is buy a new answer.

A living map of modern AI — kept current every morning