Home › Use Cases › Recoverable expense coding
Use caseUC0502
🧪 Use-case kit · runnable

Recoverable expense coding

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A property expense entry is posted to a general ledger account and somebody has to say what the lease lets the landlord recover. The managing agent has the entry's header, the recovery schedule's clauses on file, the invoice narrative in prose and a notes block the desks have written to each other — and has to place it on one of eight rungs and name the clause that puts it there. Nothing on the entry states a class, 42 of these 64 are miscoded on file, and the answer moves EVERY tenant's share, not one line of one statement. reading one property expense entry end to end — every clause tested for schedule, effective date, lookback and FINAL status, every note read for a strike and joined to the dated clause it names, and the coding basis resolved against the three printed terms — before the eight-rung ladder is applied and the queue chosen.

Audience

A managing agent's property accounting or lease administration desk coding recoverable expenses for a recovery year, and the manager deciding whether a model call is worth buying for it. This report's own answer for this corpus is NO — the call is beaten significantly by a free rule that reads the GL account's own name — and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual property expense entries

The corpus is 64 property expense entries, 0.13 MB (txt 64). A real entry names a real property and a real landlord's general-ledger detail, and the correct answer depends on an individually negotiated, confidential lease — so there is no public corpus of (expense entry, correct recovery class) pairs and there cannot be one. Generated at a fixed seed with the ambiguity planted deliberately: the narrative disagreeing with the GL account's name, capital dressed as repair, a strike note dated before the clause it names, a near-miss schedule id, and a cost above the schedule's own capital threshold. 32 fact patterns x 2 mirror twins, and the twins are built so a rule that does not READ cannot tell the halves apart — which is exactly what twin_same measures.

The corpus

  • The 64 property expense entriesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your property expense entries. That is the whole change — there is no database to migrate.

One property expense entrie, as the model receives itPEX-0001.txt · 1 of 64
PROPERTY EXPENSE ENTRY  PEX-0001
REGISTER EXTRACTED  2026-01-05   STANDARD  LRS-2026
PROPERTY  PRP-236   RECOVERY YEAR  2025   SCHEDULE  RSV-673
GL ACCOUNT  GLA-2259   GL TERM  drainage-works   CODED ON FILE  pool-controllable
VENDOR  VND-6577   WORK ORDER  WO-18444   POSTED  2025-09-21
PROJECT  PRJ-7476   PROJECT TOTAL  5095.82   ENTRY AMOUNT  5095.82
YEAR START  2025-01-01   AS OF  2025-12-31
LOOKBACK DAYS  540   CAPITAL THRESHOLD  25000.00   SERVICE LIFE YEARS  5

[1] SCHEDULE CLAUSES ON FILE
ID          KIND           SCHEDULE   EFFECTIVE   STATUS  SUBJECT                  EFFECT
CLS-30001   SCOPE-MAP      RSV-673    2025-04-29  FINAL   common-lobby             SCP-293
CLS-30002   POOL-ASSIGN    RSV-673    2024-09-09  FINAL   SCP-910                  controllable
CLS-30003   CODING-BASIS   RSV-673    2025-06-26  FINAL   -                        work-performed
CLS-30004   SCOPE-MAP      RSV-673    2024-07-11  FINAL   drainage-works           SCP-910
CLS-30005   CAPITAL-RULE   RSV-673    2025-04-05  FINAL   -                        capitalised
CLS-30006   SCOPE-MAP      RSV-673    2023-06-11  FINAL   retail-mall-run          SCP-491
CLS-30007   POOL-ASSIGN    RSV-673    2023-08-21  FINAL   SCP-491                  controllable
CLS-30008   POOL-ASSIGN    RSV-673    2024-09-18  FINAL   SCP-293                  uncontrollable

[2] INVOICE NARRATIVE
Vendor statement: serving the drainage-works - retail-mall-run, on the common-lobby; service life stated as 4 years.

[3] NOTES
Portfolio desk: the finding of 2024-06-11 that CLS-30004 was rescinded for RSV-673 was itself upheld; CLS-30004 no longer applies and nothing replaced it.
Property accounting: the declared lookback, capital threshold and service-life threshold are reprinted on the scope line above.

Abridged — the file continues.

The outcomeWhat a good result looks like

One entry in, one coded entry out: the recovery class, one of the eight rungs of LRS-2026, and the exact set of qualifying clause ids; then — in code, for every arm alike — the pool, the queue, the recode suffix, the route, the amount under test and the totals in front of the desk. It never recodes the entry, never touches the chart of accounts and never alters a tenant's share.

And when it cannot

And when it cannot: undetermined, with one short unresolved line saying what is missing, and no class named. 10 of the 64 entries are undetermined in the key. ⚠︎ A fixed undetermined reply reaches all 10 at 15.62% precision, so the paid call's 3 hits of 16 answered (18.75% precision, 30.0% recall) may never be quoted as a result without both halves — and free code reaches 66.67% precision on the same rung.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your schedules carry a strike as a STRUCTURED COLUMN — a struck-by id or a status — rather than as a sentence in a notes block. — the free GL-account rule, and do not buy a call at all
    32 of 64 classes and 20 of 64 whole rows for $0.00, against the call's 22 and 12. On the 34 entries no note decides it is 20 of 34 where the call is 8, p = 0.0075.
  • Your desk notes really do decide the class, and you want the class rather than the queue. — measure a constant and a word list on YOUR corpus first — both are free and both are in this kit
    on the 30 entries here where a note strikes a clause, the paid call takes 4 and a fixed undetermined reply takes 4: p = 1.0000. The reading win is worth nothing on this corpus, and the kit was built to test exactly that claim.
  • You need the pool, the queue, the recode flag or the route. — src/rules.py::rollup, for nothing — but read the next line before you assume it is free
    every one of those is a lookup ON the recovery class, so all four move with the reading. They are NOT free arithmetic: the 768-rule sweep ceiling is route 40, queue 44, recode 50 of 64, and no arm is saturated on any of them.
  • You are coding entries where the GL term and the schedule's coding basis name DIFFERENT terms. — the free last-term rule, and treat the call as a second opinion at most
    on the 34 gl_term_not_basis entries the call is 6 whole against last-term's 13; on the 30 where they agree it is 6 against gl-account's 16.
  • You want a defensible number for a buying decision on your own corpus. — clone, replace data/corpus and data/gold.jsonl, run evals.baseline --all --sweep and evals.check_labels, then decide
    every floor, every grader, the 768-rule sweep and every paired test are pure code and cost $0.00, so the comparison is reproducible before a single call is bought.

At a glanceHow the whole thing runs

22%rechecked recovery class correct
1,395 msp50, end to end
$0.47per 1,000 recovery classifications · the cheapest published long-context card

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own entries in the same plain-text shape — header, clauses on file, invoice narrative, notes block — and write data/gold.jsonl with one {entry_id, recovery_class, cited} per entry. EVERY QUALITY FIGURE ON THIS PAGE STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per entry for a reading a column already gives you. That is the case against the best-fitting scenario (“Your schedules carry a strike as a STRUCTURED COLUMN — a struck-by id or a status — rather than as a sentence in a notes block.”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?AN ENTRY WHOSE STRIKE IS A STRUCTURED COLUMN rather than a sentence in a notes block. The whole contest here is reading notes; on the 34 entries where no note decides anything, free code that reads no note is 20 of 34 whole and the call is 8 — p = 0.0075 AGAINST. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE CALL IS WORSE THAN FREE CODE, OR ONLY NOT BETTER. On the class it is worse and the test says so: 22 against 32, p = 0.0309 AGAINST. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-pool-coding. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board and scores all ten free arms offline, and the scored run replays from the committed results/ rather than being re-bought; --selfcheck re-derives 704 of 704 arm-entry cells and exits 0. ⛔ MEASURED ON A SCRATCH COPY WITH results/cache-*.jsonl DELETED, WHICH IS WHY THE REPLY CACHES SHIP: the board falls from 704 to 654 arm-entry cells and the PAID arm reads recovery_class null, draft no_reading and unparsed true on EVERY entry — a free arm recomputes for nothing, a bought one cannot. results/*.jsonl is deliberately NOT in .gitignore and the reason is written in the file.

A living map of modern AI — kept current every morning