The business caseThe problem this solves
A lease abstract is a structured record of a lease's commercial terms — rent, dates, area, options, notice periods — typed once out of a document pack. A portfolio is then administered from it for years while the pack keeps growing: amendments, drafts that were never signed, amendments that were terminated. Nobody re-reads the pack. The record is what feeds the rent roll, the audit confirmation and the lender's estoppel certificate, and the only way to know whether it is still true is to open every document and check every field by hand. The read-through: one person opening a document pack, finding the clause that governs each field of a record somebody else typed, deciding whether it still says the same thing, and — the step that actually gets skipped — working out whether a value that disagrees was ever right. It does not replace the correction. Nothing here writes a value into an abstract, proposes a replacement, approves one or names who may.
Audience
A lease administrator working a records-quality cycle, and the portfolio manager who signs the rent roll off. The decision this report is for is narrower than it looks: not 'is our data good' but 'which of these fields can I still evidence from the pack, and where exactly do I look'. The answer to four fifths of it is a regular expression, and this report says which fifth is not. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual lease abstract audit packs
The corpus is 62 lease abstract audit packs, 0.22 MB (txt 62). It is generated because it has to be. A real lease abstract audit pack is a joined extract of somebody's document management system and their lease administration record: an executed lease, its amendments, the drafts that were never signed, and the eighteen fields a portfolio is run from. Every page of it is commercially confidential and most of it is somebody's signature. ⚑ AND A GENERATED CORPUS BUYS SOMETHING A FETCHED ONE CANNOT HERE: the answer key needs to know, for each field, which documents state it, with what value, and which of those the register holds in force. On a real pack that is a judgement; on a generated one it is the construction, so the key is DERIVED rather than written and a second implementation can check it.
The corpus
- The 62 lease abstract audit packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — every one of the 62 packs, the document register and the whole answer key are generated in-process from seed 20260902, and
tools/build_corpus.py --checkfails if a single byte moved. Verified byte-identical under two PYTHONHASHSEEDs.
Swap this folder for your own material and the kit is pointed at your lease abstract audit packs. That is the whole change — there is no database to migrate.
LEASE ABSTRACT AUDIT PACK LAB-0001
Prepared 2026-09-02 under LAA-2026 | portfolio Marrowgate Estates
PACKET FACTS
Property Calderwell Point, Suite 100
Tenant Thrayle Analytics Ltd
Landlord entity Marrowgate Point Holdings LLC
Abstract record ABS-4100, last touched 2023-01-01
Pack reference LAB-0001
DOCUMENT REGISTER
doc title status executed
L1 Lease Agreement EXECUTED 2016-12-01
THE ABSTRACT ON FILE
field value on file
Base rent (monthly) $18,600.00
Commencement date 2017-02-01
Expiration date 2027-01-31
Rentable area 22,337 sq ft
Renewal option Two 3-year options
Permitted use General office and administrative use
Parking spaces 45 spaces
Holdover 175% of base rent
THE DOCUMENT PACK
L1-1.3 Rentable Area. The Premises are agreed for all purposes of this Lease to contain Twenty-Two Thousand Three Hundred Thirty-Seven (22,337) rentable square feet.
L1-2.2 Expiration Date. The Term shall expire at 11:59 p.m. on January 31, 2027 unless sooner terminated in accordance with this Lease.
L1-4.1 Base Rent. Tenant shall pay to Landlord base rent of Eighteen Thousand Six Hundred Dollars ($18,600.00) per month, payable in advance on the first day of each calendar month.
L1-7.1 Permitted Use. The Premises shall be used solely for general office and administrative use and for no other purpose without Landlord's prior written consent.Abridged — the file continues.
The outcomeWhat a good result looks like
One row per audited field: a verdict from four (SUPPORTED, STALE, CONTRADICTED, UNSUPPORTED), the clause id it was read from, whether that clause states the value on file, whether an earlier executed document did, the clause quoted verbatim, a confidence and one sentence of basis. 495 of 496 rows correct on the scored run (99.8 pct) with the clause quoted right on 495 of them, and 61 of 62 packs entirely correct on both graded fields.
And when it cannot
And what it does when it cannot. ONE row of 496 was wrong: LAB-0010's security deposit, where the pack carries a letter of credit delivered IN LIEU OF a cash deposit in the identical amount. The call named that clause and answered CONTRADICTED where the key's LA-8 says UNSUPPORTED — a defensible reading of a rulebook that does not say whether a clause about the topic is a clause about the field, at the lowest confidence it gave any wrong answer (0.86 against a 0.99 median). ⚠︎ AND THE LARGER HONEST ANSWER IS ABOUT THE CORPUS: at 99.8 pct this labelled set can no longer separate the paid arm from a perfect one. One row of headroom is not a measurement, and the next purchase is a harder set rather than a bigger model.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your document packs have clauses headed with the field's own name, and your abstract holds values in the same units the clauses state them in — the free rules floor alone — evals/baseline.py
It scores 82.5 pct for $0.00, ties the paid call on 10 of the 14 construction cases, and gets every draft and terminated amendment right because the pack prints its own register. Four fifths of this job is a heading map. - Your packs file clauses under headings that do not carry the field's name, or your abstract holds a unit the lease does not state — the paid call
That is the whole measured margin: 87 rows in four families — 28 where the clause is headed something else, 24 where the value is stated per annum or in days, 19 where a look-alike instrument carries the same figure, 16 where the SUPERSEDED statement is the prose one. The floor scores 0 of 87 on those and the call scores 86. - You need STALE told apart from CONTRADICTED — a maintenance failure from an abstracting one — the paid call, and read
also_in_prioron its own
130 of 130 against the floor's 114. It is the question the whole audit is for: one answer sends somebody to find the amendment that was never carried through, the other says the value was never right and the fault is in the abstracting.
And where nothing here is good enough:
- You want a number for how good this will be on YOUR packs — nothing here — run your own corpus through evals/check_labels.py first
99.8 pct is a claim about these 496 generated rows, not about lease abstracts. The corpus is near-saturated for the paid arm and has one row of headroom left.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own packs (five panels, a clause id at the head of every clause line), data/packs.json with your own register rows (one per pack, each document carrying status from EXECUTED / DRAFT / TERMINATED / WITHDRAWN and an order among the executed ones), and data/gold.jsonl with your own key — one row per printed abstract field, carrying the governing clause_ref, the two reading booleans, the verdict and the citation with its character span. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for the four fifths. It is most of the work and none of the value. That is the case against the best-fitting scenario (“Your document packs have clauses headed with the field's own name, and your abstract holds values in the same units the clauses state them in”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A scanned or photographed document pack. Every arm here — the model's reading, the floor's heading map and the independent key check — rests on a printed, panelled text file with a clause id at the head of every line. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A second run of the same corpus. There is one paid run and one adversarial arm; run-to-run variance on this workload is unknown. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-02 — r001-abstract-audit. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured rebuilds all 62 packs from seed 20260902, runs the independent label gate to 0 disagreements, scores both free floors and runs the stub end to end — the whole path in 1.1 seconds, with no network and no pip install. The board renders on 127.0.0.1:9284 with every pack, both floors, LAA-2026 and the committed run replayed from its own result file; only 'Check with the model' needs a key, and it is disabled and says so.



