The business caseThe problem this solves
A subcontractor bills a professional-services firm for a month of consultants' hours, and before AP can match the invoice it has to agree with three other records: the approved timesheets, the rate card the subcontract actually carries, and what is left under the subcontract's not-to-exceed. The AP desk's three-way match reads those records' COLUMNS — a STATUS printed on an export, a rate on a card, an invoiced-to-date figure. What decides the answer is often a sentence: an approval the engagement desk reopened after the export, hours recoded to another engagement, an amendment rate withdrawn before it was countersigned, a corrected entry whose original was already billed. The desk's own status line disagrees with the procedure on 39 of the 64 invoice packs in this corpus. Opening one invoice pack, checking every line against the timesheet entry it cites and that entry's approval at the AS AT date, reading each note to see whether it reopens, recodes, withdraws or resubmits something on this engagement, pricing each line at the rate in force for the entry's role, and testing what is left against the subcontract ceiling.
Audience
The accounts-payable desk and the subcontract desk at a professional-services firm, working a period's subcontractor invoices before the payment run — and the engagement desk that owns the timesheet queries an exception raises. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual invoice packs
The corpus is 64 invoice packs, 0.26 MB (txt 64). It is generated because it has to be. A real subcontractor invoice book is a firm's own commercial record: named consultants on every timesheet, client engagements, negotiated rates that are confidential, and engagement-desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 invoice packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every subcontractor is an invented code and trading name, every subcontract, engagement, timesheet entry, rate row and invoice is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a consultant is a resource code and a note speaks for the engagement desk, the subcontract desk, the AP desk or a vendor remittance note. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your invoice packs. That is the whole change — there is no database to migrate.
================================================================================
SUBCONTRACTOR INVOICE VERIFICATION -- ONE INVOICE, ONE AS-AT DATE
================================================================================
FILE SCI-0001
SUBCONTRACTOR V-5101 Brackenfold Analytics
SUBCONTRACT SC-7300
ENGAGEMENT ENG-3100 client CL-2200, a regional grocer
INVOICE SI-26100
INVOICE DATE 2026-07-01
BILLING PERIOD 2026-06-01 to 2026-06-30
AS AT 2026-07-04
PACK COMPILED 2026-07-08
PAYMENT TERMS net 30 days from receipt
CURRENCY USD
-- CONTRACTED RATE CARD (schedule B and every amendment printed on it) ---------
ROW ROLE RATE/HR EFFECTIVE SOURCE
RC-7300-1 PRINCIPAL CONSULTANT 240.00 2026-01-01 SC-7300 schedule B
RC-7300-2 SENIOR CONSULTANT 185.00 2026-01-01 SC-7300 schedule B
RC-7300-3 CONSULTANT 140.00 2026-01-01 SC-7300 schedule B
RC-7300-4 ANALYST 95.00 2026-01-01 SC-7300 schedule B
-- SUBCONTRACT CEILING ---------------------------------------------------------
NOT-TO-EXCEED 90,000.00
INVOICED TO DATE 79,235.00
REMAINING BEFORE THIS INVOICE 10,765.00
-- TIMESHEET ENTRIES ALREADY INVOICED ON THIS SUBCONTRACT ----------------------
INVOICE DATED ENTRY RESOURCE WORK DATE HOURS
SI-25100 2026-06-01 TS-40000 R-4400 2026-05-26 7.50
SI-25100 2026-06-01 TS-40001 R-4401 2026-05-27 8.00
-- TIMESHEET EXTRACT (exported 2026-07-02 from the timesheet system) -----------
ENTRY WORK DATE RESOURCE ROLE ENGAGEMENT HOURS STATUS
TS-50006 2026-06-01 R-4400 SENIOR CONSULTANT ENG-3100 7.50 APPROVEDAbridged — the file continues.
The outcomeWhat a good result looks like
One invoice pack in, one row out: which cited entries are approved and chargeable at AS AT, which an earlier invoice already billed, which rate rows are in force, every line's flags, the supported amount to the cent, every exception and one of five SIV-2026 verdicts. 57 of 64 packs come back with all six graded fields right, against 40 for the best free floor and 0 for the AP desk's own status line.
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 7 packs it got wrong are named in the kit README with what it answered, and every one is the entries reading: 4 late approvals it left out, 1 recoded entry it kept — the one invoice in 64 it passed as SUPPORTED — and 2 approved entries no note mentions that it dropped. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your timesheet system records a reopened approval, a recoding and a withdrawn amendment as a STATUS, and your AP export lists every entry already billed — the free modal floor, and do not buy a call at all
40 of 64 packs for $0.00. Duplicate lines, prior billing, hours over, off-engagement entries, rate creep, role uplifts and the ceiling are all decidable from columns and dates, and on those 37 packs the floor gets 37. - Approval changes, recodings and withdrawn amendments arrive as free text — an engagement-desk note, an email pasted into the pack — the paid call
This is the whole product. On the 27 packs where a sentence decides a reading the paid arm is 21, the modal floor 3 and the best vocabulary floor 13. - You want unapproved hours kept OUT of the payment run above all — the paid call, and read the vocabulary floor beside it
The modal floor passes 19 of the 51 flagged invoices as SUPPORTED; the paid arm passes 1 (SCI-0029). The one it lets through is a recoded entry, and the vocabulary floor catches recoded entries 5 of 5. - You want the AP desk's own three-way match audited — either paid or free — both beat it comprehensively
The desk's status line disagrees with SIV-2026 on 39 of 64 packs and gets 0 whole: it returns no reading at all. It is published as an arm so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Approvals routinely come through after the timesheet export — neither alone — re-export the extract at AS AT
The paid arm TIES the modal floor at 2 of 6 late approvals: it never admitted an approval dated before AS AT. A fresh extract makes the reading a column again. - Your subcontracts carry retainage, expenses, fixed-fee milestones or a rounding tolerance — neither, yet
No file in this corpus does. The unit of work is time-and-materials lines citing timesheet entries, and every percentage on this page is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own invoice packs in the same six-block shape and data/invoices.json with your own subcontract register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per invoice for arithmetic you already have. That is the case against the best-fitting scenario (“Your timesheet system records a reopened approval, a recoding and a withdrawn amendment as a STATUS, and your AP export lists every entry already billed”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An invoice line that cites no timesheet entry — a fixed-fee milestone, an expense, a lump sum. The whole reduction is the line-to-entry join; a line with nothing to join to reads as unapproved hours. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-subcon-invoice. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.








