The business caseThe problem this solves
A tenant pays a security deposit at signing. Then the lease is renewed and the deposit raised, a pet deposit is added, a top-up bounces and is paid again, a year's interest is keyed twice after a batch fails to confirm, and at move-out a deduction is posted and backed out. By the cut-off there are three records of what this deposit is: what the lease and its amendments require, what the ledger says moved, and the interest rule the operator recorded for the property. They disagree, and the disagreements are the kind that end in a dispute — a raise nobody collected, a deposit held above what the lease allows, a deduction nobody with authority approved. The ledger system's own status line is computed from the columns and is wrong on 38 of the 62 lease files in this corpus — on 20 of them because the thing that decides the answer is a sentence in the notes, and on the other 18 the columns alone decide it and free code gets every one right. Opening one lease's deposit file, reading every amendment against the deposit at signing, reading every ledger posting against the lease, reading each note to see whether it is about this lease and whether it takes a row out, recomputing the interest under the property's recorded rule, and checking every deduction and refund for an approver code and a permitted category.
Audience
The property-management accounts desk working a period's deposit exceptions, and the property manager who has to decide a move-out disposition once the evidence is in front of them — this kit hands them the evidence and never the decision. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual lease files
The corpus is 62 lease files, 0.17 MB (txt 62). It is generated because it has to be. A real deposit ledger is a landlord's own client-money record: tenant names, unit addresses, bank references and move-out disputes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that takes a row out — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies. The interest rules are invented too: the corpus DECLARES a recorded rule per property, or declares that none is recorded, and never quotes a jurisdiction.
The corpus
- The 62 lease filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every property is an invented code and name, every lease, tenant account, amendment, posting and amount is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note is from the accounts desk or the leasing office, and an approver is a code. evals/check_labels.py sweeps all 62 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your lease files. That is the whole change — there is no database to migrate.
==============================================================================
SECURITY DEPOSIT LEDGER RECONCILIATION -- ONE LEASE, ONE CUT-OFF
==============================================================================
FILE DL-0001
PROPERTY P-101 Larch Field Apartments
UNIT Unit 1A
LEASE LSE-50100
TENANT ACCOUNT TA-70100
LEASE START 2025-01-01
LEASE END 2029-12-31
MOVE-OUT --
ASSIGNED --
CUT-OFF 2026-03-31
TOLERANCE 1.00
CURRENCY USD
-- THE LEASE DEPOSIT TERMS ---------------------------------------------------
DEPOSIT REQUIRED AT SIGNING 1,800.00
DEPOSIT TRANSFERS ON ASSIGNMENT YES
DEDUCTIONS THE LEASE PERMITS CLEANING, DAMAGE, KEYS, UNPAID-RENT
RECORDED INTEREST RULE IR-2 simple interest, 1.25 pct of the principal held, per full lease year, posted on each anniversary
INTEREST RULE SOURCE the property's recorded rule register; no jurisdiction rule is computed here
-- DEPOSIT AMENDMENTS AGAINST THE LEASE --------------------------------------
AMEND SIGNED EFFECTIVE CHANGE DELTA STATUS
AMD-2201 2025-12-12 2026-01-01 renewal letter, deposit raised +300.00 EXECUTED
AMD-2202 2025-12-20 2026-01-01 pet deposit added +200.00 EXECUTED
-- DEPOSIT LEDGER POSTINGS ---------------------------------------------------
POSTING DATE LEASE TYPE AMOUNT CATEGORY APPROVER STATUS REFERENCE
DLP-81001 2024-12-28 LSE-50100 RECEIPT +1,800.00 -- -- POSTED DEPOSIT AT SIGNING, CHEQUE 400100
DLP-81010 2026-01-01 LSE-50100 INTEREST +22.50 -- -- POSTED INTEREST POSTING IR-2Abridged — the file continues.
The outcomeWhat a good result looks like
One lease file in, one row out: which amendments are in force at the cut-off, which postings stand with each row quoted verbatim, the required deposit, the principal held, the interest due under the RECORDED rule (or no figure where none is recorded), the balance held, and one of seven SDR-2026 verdicts. 59 of 62 lease files come back completely right, against 51 for the best free code this kit could write and 0 for the ledger system's own status block.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, nothing stopped at the ceiling and no call failed. The 3 lease files it got wrong are named in the kit README with what it answered, and all 3 are notes that name no row id: DL-0008 kept a renewal letter a note says was replaced in its entirety, DL-0015 kept an amendment a note says was cancelled before it took effect, and DL-0020 counted a top-up that came back unpaid beside the payment that replaced it. On DL-0015 the best free code is right and the paid call is not. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your amendment register marks a replaced or rescinded amendment in a column, and your ledger voids its own bounced, re-keyed and backed-out postings — the free modal floor, and do not buy a call at all
38 of 38 column-decided lease files for $0.00. Cancelled amendments, VOID and MEMO postings, other-lease postings, the cut-off, approver codes, deduction categories and the recorded rule are all decidable from columns and registers, and the floor gets every one of them. - Replacements, rescissions, bounced payments and re-keys arrive as free text — a leasing-office note, an accounts-desk comment, a line pasted from the bank — the paid call
This is the whole product. On the 24 lease files where a sentence decides the reading the paid arm is 21, the best free code 13 and the modal floor 0. - Your notes describe the row by amount and date rather than by id, and rescissions are the common case — the free rules floor, and read the paid arm beside it
On rescinded_amendment the best free code is 5 of 5 and the paid call 4 of 5 — DL-0015 is the one file in the corpus where free code beats it. All three of the paid arm's misses are id-less notes. - A property has no interest rule recorded, and you need that stated rather than defaulted — either paid or free — the station decides it
INTEREST-UNVERIFIABLE is 5 of 5 on every arm that reads the file, because src/policy.py computes no figure where data/interest_rules.json has no row. The ledger system prints 0.00 on all five. - You want the ledger system's own status line audited — either paid or free — both beat it comprehensively
The status block disagrees with SDR-2026 on 38 of 62 lease files, and it is published as an arm precisely so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Your portfolio needs a disposition DECIDED — what the tenant is owed at move-out — neither, and by design
This kit has no disposition authority. It names an unsupported deduction and why; the decision is a named property manager's act, and the answer contract has no field to hold it.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own lease files in the same shape, data/leases.json with your own register and data/interest_rules.json with the rules you have actually recorded, then run python3 -m evals.run --run-id b000-<yours>-rules --floor rules — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per lease file for arithmetic you already have. That is the case against the best-fitting scenario (“Your amendment register marks a replaced or rescinded amendment in a column, and your ledger voids its own bounced, re-keyed and backed-out postings”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A ledger export with no STATUS column. R-2's VOID and the whole distinction between a posting printed and a posting that stands start from it. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-deposit-ledger. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run, and python3 -m src.app --check re-derives every station verdict for $0.00. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.







