The business caseThe problem this solves
A subcontractor bills a general contractor every month, and before the project manager certifies anything the application has to agree with four other records: the schedule of values the subcontract carries at the AS AT date, the progress measurement that stands for each item, the delivery tickets that evidence stored material under the subcontract's own stored-materials clause, and the previous certificate. Project accounting's pre-check reads those records' COLUMNS — a change order's printed status, the latest measurement row, every ticket on the log. What decides the answer is often a sentence: a change order the contracts desk executed after the log was printed, a re-measure field engineering voided, a load the materials desk returned or saw built in. The pre-check's own status disagrees with the procedure on 38 of the 64 pay application packs in this corpus. Opening one pay application pack, deciding which change orders are really in force at AS AT, which progress measurement stands for each item and which delivery tickets still evidence stored material, reading every note for an execution, a rescission, a voided re-measure or a returned load and its date, then re-adding the continuation sheet and testing each line against its scheduled value, its measurement and its evidence.
Audience
The project-controls desk and project accounting at a general contractor, working a period's subcontractor pay applications before the project manager certifies — and the contracts, field engineering and materials desks that own the queries an exception raises. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual pay applications
The corpus is 64 pay applications, 0.30 MB (txt 64). It is generated because it has to be. A real subcontractor pay application book is a contractor's own commercial record: named project managers certifying, named foremen and field engineers on every measurement, negotiated schedules of values and change-order prices that are confidential, and desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 pay applicationsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every subcontractor is an invented vendor code and trading name, every project, subcontract, change order, measurement and delivery ticket is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for the contracts desk, field engineering, the materials desk, project accounting or a subcontractor cover note. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your pay applications. That is the whole change — there is no database to migrate.
================================================================================================
SUBCONTRACTOR PAY APPLICATION CHECK -- ONE APPLICATION, ONE AS-AT DATE
================================================================================================
FILE SPA-0001
SUBCONTRACTOR V-6101 Quarlesby Concrete Works
TRADE concrete
SUBCONTRACT SC-8100 dated 2026-01-12
PROJECT PJ-4100 a three-storey medical office building
APPLICATION PA-8100-03 (application no. 3)
PERIOD 2026-06-01 to 2026-06-30
AS AT 2026-06-30
APPLICATION DATE 2026-07-02
PACK COMPILED 2026-07-06
STORED MATERIALS ON SITE ONLY (subcontract article 7)
CURRENCY USD
-- SUBCONTRACT SCHEDULE OF VALUES (base schedule and every change order on the log) ------------
ROW ITEM DESCRIPTION VALUE SOURCE STATUS DATED
SV-8100-01 01 Mobilization and general conditions 18,500.00 SCHEDULE EXECUTED 2026-01-12
SV-8100-02 02 Footings and foundations 158,900.00 SCHEDULE EXECUTED 2026-01-12
SV-8100-03 03 Slab on grade 176,400.00 SCHEDULE EXECUTED 2026-01-12
SV-8100-04 04 Elevated decks 312,000.00 SCHEDULE EXECUTED 2026-01-12
SV-8100-05 05 Columns and walls 134,600.00 SCHEDULE EXECUTED 2026-01-12
-- PREVIOUS CERTIFICATE (application no. 2, certified 2026-06-06) ------------------------------
ITEM WORK COMPLETED STORED MATERIALS
01 18,500.00 0.00
02 47,670.00 8,600.00
03 70,560.00 0.00
04 171,600.00 0.00Abridged — the file continues.
The outcomeWhat a good result looks like
One pay application pack in, one row out: which schedule-of-values rows are in force at AS AT, the one progress row that stands for each item, which delivery tickets evidence stored material, every line's flags, the supportable amount to date to the cent, every exception and one of five PAC-2026 verdicts. 62 of 64 applications come back with all six graded fields right, against 41 for the best free floor and 0 for project accounting's own pre-check status.
And when it cannot
And what it does when it cannot. On the published run 64 of 64 replies parsed, 0 stopped at the 1,000-token ceiling, and the streamed reply stop closed 1 call after its complete answer. The 2 applications it got wrong are named in the kit README — SPA-0018 and SPA-0033, a change order withdrawn and a re-measure voided by notes dated AFTER AS AT, both followed anyway, both false exceptions on NO-EXCEPTION applications. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your change-order log, progress system and delivery log record executions, voided re-measures, returns and installations as a STATUS, dated — the free modal floor, and do not buy a call at all
41 of 64 applications for $0.00. Arithmetic breaks, lines over their scheduled value, work ahead of the measurement, stored material with no ticket and the stored-materials clause are all decidable from columns and dates, and on those 36 packs the floor gets 36. - Change orders, re-measures and deliveries are changed in free text — a contracts-desk note, a field email pasted into the pack — the paid call
This is the whole product. On the 28 packs where a sentence decides a reading the paid arm is 26, the modal floor 5 and the best vocabulary floor 19. - Your notes are routinely written days after the AS AT date, about things that happened after it — the modal floor beside the paid call, and read the dates yourself
Both of the paid arm's misses are a note dated AFTER AS AT that it followed anyway, and the modal floor, which reads no note, gets all 5 date traps. A date comparison in code would have caught both. - You want every unsupported dollar kept OUT of the certificate above all — the paid call
On the 50 applications the key flags, the station never came back NO-EXCEPTION on the published run; the modal floor let 19 through. Both paid misses are in the cheaper direction — a query about value the evidence supports. - You want project accounting's own pre-check audited — either paid or free — both beat it comprehensively
The pre-check's status disagrees with PAC-2026 on 38 of 64 packs and gets 0 whole: it returns no reading at all. It is published as an arm so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Your subcontracts carry retainage, deductive change orders, partial draws from storage or a tolerance on measured progress — neither, yet
No file in this corpus does. The unit of work is one application, one period, one line per schedule item, and every percentage on this page is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own pay application packs in the same nine-block shape and data/payapps.json with your own register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per application for arithmetic you already have. That is the case against the best-fitting scenario (“Your change-order log, progress system and delivery log record executions, voided re-measures, returns and installations as a STATUS, dated”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A pay application whose continuation lines do not map one-to-one onto schedule items. The station sums schedule rows, measurements and tickets PER ITEM; a line covering two items, or an item split across lines, has nothing to join to. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | ONE PUBLISHED SCORED RUN, AND ITS PREDECESSOR IS NOT A REPEAT. r001-payapp-check scored 60 of 64 with the same prompt, station and ceiling but no reply stop; r002 scored 62. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r002-payapp-check. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.









