Home › Use Cases › Check every subcontracted client line against the markup terms that govern it
Use caseUC0501
🧪 Use-case kit · runnable

Check every subcontracted client line against the markup terms that govern it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A consulting firm subs specialist work out and re-bills it to the client at the markup the engagement letter allows. The cost keeps moving after the invoice goes out — a credit note from the sub, a short-paid line, an amount the firm absorbed, a corrected entry recorded as a note. Today a billing-desk analyst opens the pack, reads every adjustment sentence to decide which are applications of THIS line's own cost record and when, reads the markup terms to decide which one is actually in scope, and then does the arithmetic twice — as the line was billed, and as it stands today. The part everyone gets wrong is the sentence that names an amount and an adjustment reference and is not an application at all. The sentence-by-sentence read of one cost record's adjustments and one clause set of markup terms, twice over two dates — not the decision to credit the client, not the markup change, and not what the subcontractor is owed, which this kit never computes.

Audience

The billing desk of a professional-services firm and the partner who signs off a client credit or a markup exception. The decision is narrow: which subcontracted client lines need a partner to look again, and why. This report answers it with a flat NO on whether the call is worth making — 31 packs against free code's 30, p = 1.0 — and a qualified YES on three named term-scope families, which is 13 packs of the 64. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual subcontracted billing lines

The corpus is 64 subcontracted billing lines, 0.14 MB (txt 64). Synthetic, and said to be, because no firm's subcontract clause sets and client invoice lines can be published. The material is plain text in one block layout for the same reason the reading is the product: the whole question is which SENTENCE decides, so a corpus of tables would measure a different use case. Two earlier builds were measured and thrown away for saturating, and the current one exists to make three things hard that a column cannot express — an adjustment that names an amount and is not an application, a scope clause naming two ids of the same kind, and a term taking its scope from another term by reference.

The corpus

  • The 64 subcontracted billing linesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator cannot do — including the TWO builds that were measured and thrown away for saturating (the first scored governing_term 64 of 64 to a free arm that pulled three fields out of each term row; the second put SETUP-STALE on 36 of 64 packs because every hard-to-read term treatment also billed a wrong rate), and the vocab_coverage assertion that now fails the BUILD on any rulebook value no item carries.

Swap this folder for your own material and the kit is pointed at your subcontracted billing lines. That is the whole change — there is no database to migrate.

One subcontracted billing line, as the model receives itPTL-0001.txt · 1 of 64
PASS-THROUGH MARKUP PACK  PTL-0001
PACK ASSEMBLED  2026-02-19   STANDARD  PTM-2026

ENGAGEMENT  ENG-4401   CLIENT  CLI-7101   SUBCONTRACTOR  SUB-2201

[1] CLIENT BILLING LINE
FIELD                VALUE
LINE                 PTL-0001
INVOICE              INV-9301
INVOICED             2026-01-25
SERVICE PERIOD       2026-01
SERVICE DATE         2026-01-05
COST CATEGORY        SPC-01
COST RECORD          SCI-2201
COST BASE BILLED     4,000.00
MARKUP BILLED        8.0 per cent
AMOUNT BILLED        4,320.00

[2] SUBCONTRACT COST RECORD
REF        RAISED BY              AMOUNT
SCI-2201   Subcontract desk       4,000.00
  ENTRY  SCI-2201 records subcontracted cost of 4,000.00 under engagement ENG-4401 for service in period 2026-01.
SCI-2202   Subcontract desk       3,000.00
  ENTRY  SCI-2202 records subcontracted cost of 3,000.00 under engagement ENG-4401 for service in period 2026-01.
SCI-2203   Subcontract desk       3,097.01
  ENTRY  SCI-2203 records subcontracted cost of 3,097.01 under engagement ENG-4401 for service in period 2026-01.
ADJ-4100   Subcontract desk       150.00
  ENTRY  ADJ-4100 adjusts the cost recorded under SCI-2201 by 150.00 in reduction; recovery is recorded to 2026-01-13; the adjustment was applied on 2026-01-29.
ADJ-4101   Billing desk           150.00
  ENTRY  ADJ-4101 adjusts the cost recorded under SCI-2202 by 150.00 in reduction and the reconciliation was run on 2026-01-08; the adjustment was applied to SCI-2202 on 2026-01-19.

[3] MARKUP TERMS ON FILE
REF        RATE            RAISED BY
MT-3300    8.0 per cent    Engagement desk
  TERM   MT-3300 records the subcontracted-cost markup as 8.0 per cent for engagement ENG-4401, effective from 2025-08-08.
MT-3301    12.0 per cent   Pricing desk

Abridged — the file continues.

The outcomeWhat a good result looks like

One subcontracted billing line in, one answer out: the settled cost as billed, the settled cost as at today, the governing markup term, the two verdicts, the cause and a 200-character desk note. A good result is all six graded cells right on a pack — 31 of 64 here, against the best free arm this kit ships at 30. It is not a good result.

And when it cannot

And what it does when it cannot. 64 of 64 replies parsed, 0 unparsed, 0 failures, 0 calls at the ceiling, 0 requests the provider never started and 0 notes over budget. There is no abstention: every pack gets six cells, so a pack it read badly comes back looking exactly like a pack it read well. 33 of 64 packs come back with at least one cell wrong and the kit cannot say which 33. ⛔ Worse, the one failure shape src/station.py cannot see is a reading that parses perfectly and is wrong: four of the 24 adversarial replies DENY IN PROSE exactly what their own number takes, and the station does perfect arithmetic on the wrong cost and returns a clean row.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know whether paying for this reading beats a careful rules engine — the floor of record (b000-passthru-markup-domain_scope), and do not buy the call
    31 packs against 30, exact McNemar 16/15, p = 1.0 — no separable difference, on 34 packs of available room. The honest answer is that the money bought nothing overall on this use case, at this tier, at these settings.
  • Your markup terms name two identifiers of the same kind in one scope clause, or take their scope by reference to another term — the scored run (r001-passthru-markup)
    this is the whole reason to pay, and it is 13 packs. Across t_pair_eng, t_pair_cat and t_ref_out the call takes 10 of 13 where free code takes 1 — the three families the corpus was rebuilt to create, and a grammar no word list reaches.
  • Most of your lines are clean — nothing moved the cost after the invoice — the floor of record, or the tables-only arm
    ⛔ the call is WORSE here. On the nine c_clean packs it takes 2 against free code's 7: it reads movement into packs that never moved. That is half of why a 34-pack headroom nets to 1.
  • You only need the settled cost as the client was billed — the tables-only arm (b000-passthru-markup-columns), at $0.00
    it is THE BAR on that cell at 48 of 64, because only 16 of these 64 packs carry an application dated before the invoice went out. The call reads 54, 13 packs to 7, p = 0.263 — not significant. Three quarters of that cell is free.
  • You need to be certain the pack never changes a markup, credits a client or moves a billing setup — any arm — the cap is a SHAPE, not a behaviour
    data/fields.json offers no field any of those acts could be written into and src/prompt.py asserts that at import. 0 breaches on every committed arm, on the 38 refusal probes the corpus already carries, and on all 24 attacked trials including the authority framing that orders three of them together.
  • You want to know what a reading is worth on your OWN corpus before spending — evals/baseline.py, at $0.00
    all seven free arms, both columns, plus --slice which scores any candidate slice across every free arm. Everything on this page except the two paid runs re-derives for nothing, so your own floor and your own per-cell bar are free to measure first.

At a glanceHow the whole thing runs

48%whole row rechecked pct
1,401 msp50, end to end
$2.57per 1,000 subcontracted billing lines · the fast tier

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same block layout, write data/lines.json with one row per pack, and produce data/gold.jsonl with one answer per pack. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Quoting 48.4% without that sentence. It is the single most misreadable number in the kit. That is the case against the best-fitting scenario (“You want to know whether paying for this reading beats a careful rules engine”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack with no [2] COST RECORD block. The settled cost is two of the three readings and there is nothing to read them off; src/policy.py raises rather than returning a basis of 0, because a 0 basis reads as a line that cost nothing rather than as a pack that could not be read. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the 48.4% whole-row figure would hold on a real firm's packs. Every figure here is measured on a generated corpus whose adjustment and clause phrasings this kit wrote, and the tuned arm reaching 384 of 384 is the proof that the phrasing is learnable once written. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-passthru-markup. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with NO key configured renders the whole board on port 9601 and scores every free arm offline: seven floors, the paired tests, the independent label check and the corpus --check all run on the Python standard library and made 0 calls. The one control that could spend, Ask the desk, renders DISABLED and says why. Measured, not asserted: the screenshot driver intercepted every request the page made and reported 0 spend attempts, and two of the 19 shots are that no-key state. ⚠︎ The reply caches SHIP for this reason — measured on a clone with them deleted, the board's pressure panel returned None for every per-trial value and printed ‘No reply denies in prose what its own number takes’ as a derived finding, which is false; the board now says so when they are missing.

A living map of modern AI — kept current every morning