Home › Use Cases › Check one expense claim against the client's own expense terms, line by line
Use caseUC0347
🧪 Use-case kit · runnable

Check one expense claim against the client's own expense terms, line by line

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A professional-services firm's people incur expenses on a client engagement and submit them one line at a time. The client's engagement letter carries EXPENSE TERMS that say which categories are reimbursable at all, what each is capped at PER UNIT, above what amount a receipt is required, which categories need advance written approval, and that a disbursement is passed through at cost. Checking a claim means joining every line to that schedule, to the engagement's own matter and dates, and to the firm's own policy where the two disagree — and then multiplying a cap by a unit count, per line, at month end, sixty claims deep. Two things in that make it a reading rather than an arithmetic exercise. The Category column is the claimant's own coding: a dinner for the firm's own people with nobody from the client there and no work done is staff entertainment however it is coded, and the client does not reimburse it. The Units column is the claimant's own count: the same charge is inside the cap for three nights and outside it for one. Get either wrong and the arithmetic is faultless and the answer is money — money queried against a colleague who did nothing wrong, or money that goes out on the claim and comes back months later as a client deduction. Reading every expense line against a nine-row expense schedule by category, against an engagement record by matter and date, and against the firm's own policy where it differs — then multiplying a cap by a unit count and comparing, line by line, by eye.

Audience

The billing or engagement-finance reviewer who sees expense claims before an invoice is assembled, and the engagement partner who owns the conversation when a line does not hold up. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual submitted expense claims

The corpus is 62 submitted expense claims, 0.19 MB (json 3 · jsonl 1 · md 2 · txt 62). Because the shape of an expense-policy failure is not the arithmetic, it is the coding. 9 of the 62 claims carry a line whose Category column is consistent with everything a comparison can check and whose description describes a different kind of expense; 7 carry a line whose Units column is consistent with everything and whose description gives a different count. Those 16 lines are the whole measurement, and the corpus is built so both directions are reachable: 8 of them the columns HOLD in error and 8 the columns LET THROUGH.

The corpus

  • The 62 submitted expense claimsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 62 claims, data/engagements.json and the whole answer key are generated in-process by the seeded builder, and the key is DERIVED by running src/policy.py over the same record the text is rendered from rather than typed.

Swap this folder for your own material and the kit is pointed at your submitted expense claims. That is the whole change — there is no database to migrate.

One submitted expense claim, as the model receives itEX-0001.txt · 1 of 62
EXPENSE CLAIM - CLIENT EXPENSE POLICY CHECK

CLAIM HEADER
  Claim              EX-0001
  Firm               Aldenmere Advisory LLP (Dublin office)
  Client             Dunhollow Retail Holdings
  Engagement         ENG-3100-01
  Claimant           M. Hallberg
  Claim period       2026-01-21 to 2026-02-17
  Submitted          2026-02-25

ENGAGEMENT AND MATTER RECORD (the firm's own record of the engagement this claim charges to)
  Engagement period        2026-01-01 to 2026-08-31
  Matter for this claim    MTR-4100
  Matter opened            2025-11-05
  Receipt threshold        Clause 7: a receipt is required for any line above $25.00
  Terms precedence         Clause 2: the client's terms are what is applied, and a cap the firm states differently is a conflict

CLIENT EXPENSE TERMS (Schedule E, terms version TE-2026-A, in force for the whole engagement)
   Category              Reimbursable  Unit                 Cap per unit  Advance approval
   Air-economy           yes           per journey          $646.00       no
   Air-premium           yes           per journey          $2,267.00     yes
   Hotel                 yes           per night            $280.00       no
   Meals                 yes           per person per day   $89.00        no
   Client-entertainment  yes           per person           $125.00       yes
   Staff-entertainment   no            -                    -             -
   Ground-transport      yes           per journey          $140.00       no
   Courier-print         yes           per item             $18.00        no
   Admin-recovery        no            -                    -             -

FIRM POLICY EXTRACT (the firm's own expense policy, printed for comparison only - Clause 2)

Abridged — the file continues.

The outcomeWhat a good result looks like

Every line of the claim carries a verdict, the clause of the client's own terms it rests on, a disposition a reviewer can act on — allowed, disallowed, waiting for evidence, or a conflict between the client's terms and the firm's own policy — the amount at issue to the cent, and one row of the claim copied verbatim as the evidence. The claim as a whole carries CLEAR or QUERY, the queried lines and the total.

And when it cannot

⚠︎ THE CLAIM-LEVEL NUMBER IS 13 OF 62 AND THAT IS NOT A TYPO. claim_all_correct requires every one of the seven graded fields on every line of the claim to be right AND the recommendation, the queried lines and the total to follow. The rulebook re-applied gets 268 of 277 line verdicts, 57 of 62 recommendations and 53 of 62 queried-line sets right — and it is capped at 13 claims by ONE behaviour: the reply quoted a row on 91 of the 235 lines that are ALLOWED, where the answer contract says null. Every free floor this kit ships scores better at claim level, and the page says so.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A firm with one client's expense terms typed up as a schedule, checking claims before an invoice is assembled — this kit
    the whole job is one claim against one schedule, and the two things a schedule cannot settle are exactly what the call is bought for.
  • A firm with no consolidated expense terms per client — consolidate the terms first
    the candidate row this kit was built from records that as an open question in its own words: client expense-terms consolidation is needed per client before checks can be complete rather than partial. There is nothing to check against until it exists.
  • Deciding who bears a disallowed line, or paying one — a person
    nothing here reimburses, writes off, absorbs or bills, and the answer contract has no field that could.
  • Checking TIME, fees or work in progress — timesheet-gap (UC0264), fee-leakage (UC0150), wip-invoice (UC0141)
    those are hours, fees and unbilled work. This kit reads expenses and only expenses.

At a glanceHow the whole thing runs

97%line verdict correct rechecked pct
3,523 msp50, end to end

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported claims and data/engagements.json at your own engagement records and expense schedules, keep the panel headings and the column order, and every free floor, the station and the whole board run unchanged with no key and no network. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Buying the call before running the free column floor. It gets 261 of 277 lines right for nothing on this corpus, and on a schedule with fewer coded-category and unit traps than this one it would get more. The paid arm is a tie against the de-memorised keyword floor at line level and BELOW every free floor at claim level. That is the case against the best-fitting scenario (“A firm with one client's expense terms typed up as a schedule, checking claims before an invoice is assembled”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A claim whose panels are not the seven this parser knows. src/rules.py splits on the headings CLAIM HEADER, ENGAGEMENT AND MATTER RECORD, CLIENT EXPENSE TERMS, FIRM POLICY EXTRACT, EXPENSE LINES, CLAIM NOTES and SIGN-OFF, and reads the line rows with one fixed-column regular expression. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates a model's variance from a real difference. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-expense-policy. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all four floors score in a few seconds and print 235, 261, 277 and 269 line verdicts. python3 -m evals.check_labels and python3 tools/build_corpus.py --check are the other two free proofs.

A living map of modern AI — kept current every morning