Home › Use Cases › Check overhead on a joint-interest bill against the agreement's own accounting procedure
Use caseUC0401
🧪 Use-case kit · runnable

Check overhead on a joint-interest bill against the agreement's own accounting procedure

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An operator bills every non-operating working-interest owner a fixed monthly OVERHEAD charge for each well the agreement covers, in place of allocating its own office and supervision cost. Two rates exist — one for a well being drilled, one for a well capable of production — and they are an order of magnitude apart. A factor is published once a year and applies from the following month. Nobody on the receiving end re-does that arithmetic every month: a joint-interest desk gets a statement, glances at the total, and pays it. The charge that is wrong is almost never the one that looks wrong — it is a well billed at the drilling rate that was already producing, a well charged for three months it was only on the account for one, or a rate computed from last year's factor, which is a real number printed on the same statement. Joining every overhead line to a well schedule by name, to two base rates by the well's actual status, to an adjustment history by billing period, and to a months count nothing on the statement prints — then doing three multiplications a line and noticing the one row whose narrative contradicts its own column.

Audience

A non-operator's joint-interest desk, reading an operator's statement before anybody decides what to pay. And the operator's own revenue-accounting team, checking its own output before it goes out. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual joint-interest overhead statements

The corpus is 60 joint-interest overhead statements, 0.18 MB (json 3 · jsonl 1 · md 2 · txt 60). Because the shape of an overhead failure is not the arithmetic, it is the STATUS and the MONTHS. 227 of the 270 lines simply reconcile, which is what a month of joint-interest statements looks like; 7 of them hide a rate computed from the factor published a year earlier, which is a real number printed on the same statement; and exactly 19 of them carry an answer that no column on the bill can reach. A corpus where half the rows were wrong would be a different job.

The corpus

  • The 60 joint-interest overhead statementsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 60 statements, data/agreements.json and the whole answer key are generated in-process by the file that also derives the key, so there is no fetch, no licence to clear and no third-party right in any of it. data/SOURCES.md states what that costs the measurement, measured rather than asserted.

Swap this folder for your own material and the kit is pointed at your joint-interest overhead statements. That is the whole change — there is no database to migrate.

One joint-interest overhead statement, as the model receives itJIB-0001.txt · 1 of 60
JOINT-INTEREST BILL - OVERHEAD RATE APPLICATION CHECK

BILL HEADER
  Statement          JIB-0001
  Agreement          JOA-3100-01
  Operator           Ternhill Petroleum
  Non-operator       Aldenmere Minerals Partners
  Unit               Calderwood Unit, Reeves County
  Billing period     2026-05 to 2026-07
  Statement date     2026-08-08
  Prepared by        M. Hallberg, for Ternhill Petroleum

ACCOUNTING PROCEDURE (Exhibit C to the agreement named in the header)
  Procedure               OHRATE-2026
  Base year               2021
  Drilling well rate      $12,300.00 per well per month, base
  Producing well rate     $1,050.00 per well per month, base
  Adjustment in force     1.0949  (published April 2026, in force 2026-05 to 2027-04)
  Working interest        0.1250  (the non-operator's share of this joint account)
  Period length           3 months

ADJUSTMENT HISTORY (the factor published each April, applied from the following May)
   Published  Factor   In force from  In force to
   2024       1.0307   2024-05        2025-04
   2025       1.0571   2025-05        2026-04
   2026       1.0949   2026-05        2027-04

WELL SCHEDULE (the wells this agreement covers)
   Well          API number      Unit                   Tract
   CALDERWOOD-1  31-113-20017    Calderwood Unit        Tract 1
   CALDERWOOD-2  32-126-20034    Calderwood Unit        Tract 2
   CALDERWOOD-3  33-139-20051    Calderwood Unit        Tract 3
   CALDERWOOD-4  34-152-20068    Calderwood Unit        Tract 4
   CALDERWOOD-5  35-165-20085    Calderwood Unit        Tract 5
   CALDERWOOD-6  36-178-20102    Calderwood Unit        Tract 6
   CALDERWOOD-7  37-191-20119    Calderwood Unit        Tract 7
   CALDERWOOD-8  38-204-20136    Calderwood Unit        Tract 8

OVERHEAD LINES

Abridged — the file continues.

The outcomeWhat a good result looks like

Every overhead line carries a verdict, the term of the accounting procedure it rests on, the amount at issue to the cent AND SIGNED, and one row copied verbatim out of the statement as evidence; the statement carries PASS or QUERY, the lines named, and the total at issue. A reviewer reads a queue instead of joining six lines to a well schedule, two base rates and an adjustment history by eye.

And when it cannot

⚠︎ THE STATEMENT-LEVEL NUMBER IS 39 OF 60 AND THAT IS NOT A TYPO. bill_all_correct requires every one of the six graded fields on every line to be right, plus the recommendation, the queried set and the total. The verdict column reads 269 of 270. What sits between them is the CITATION: the call returned a quoted row on 24 of the 227 lines that reconcile, where the answer contract says null, and one quote that is not in the statement at all. Both numbers are on every board and every run record, because a page printing either alone tells a different story.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your operators code the rate class correctly and your overhead arguments are about the factor, the well schedule and duplicates. — the free column floor alone — python3 -m evals.run --floor rules
    Every one of those is a lookup, a comparison or two multiplications. The floor is 251 of 270 line verdicts and 41 of 60 statements for $0.00 and no network.
  • Your operators' narrative column is written from a small set of standard phrases. — the keyword floor, and measure its de-memorised twin beside it
    On a corpus with a narrative pool that repeats, a word list generalises: phrase-generic keeps 16 of 23 status patterns here and still scores 262 of 270.
  • Your operators write the status and the on/off dates freely, in prose, and the money per line is four figures. — the paid call, with the station
    That is the only population this kit separates on — 19 of 270 lines, where the call is 18 and the column floor is 0, exact two-sided p = 0.000008.
  • You want the amount at issue to be right to the cent whatever the reading. — the station, on any arm
    src/recheck.py takes the arm's two readings and recomputes everything else: amounts to the cent go 224 -> 269 of 270 and amounts with the wrong sign go 6 -> 0.

At a glanceHow the whole thing runs

99.6%line verdict correct rechecked pct
2,204 msp50, end to end
$1.10per 1,000 joint-interest overhead statements · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported statements and data/agreements.json at your own agreement records. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need — and paying it once a month per statement. That is the case against the best-fitting scenario (“Your operators code the rate class correctly and your overhead arguments are about the factor, the well schedule and duplicates.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A statement whose panels are not the eight this parser knows. src/rules.py splits on BILL HEADER, ACCOUNTING PROCEDURE, ADJUSTMENT HISTORY, WELL SCHEDULE, OVERHEAD LINES, REMARKS, SIGN-OFF and END OF STATEMENT, and reads the overhead rows out of a fixed column layout. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from a real difference. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-overhead-rate. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all four floors score all 60 statements in under a second. python3 -m evals.check_labels re-derives the key's checkable half from the rendered text alone. python3 -m src.app serves the board with the committed run replayed at $0.00. SETUP, MEASURED AT TWO MINUTES: clone, copy .env.example to .env at the repo root, set PROVIDER, BASE_URL, API_KEY and MODEL. There is nothing to install — the kit is the Python standard library end to end — and none of the four commands above needs any of it.

A living map of modern AI — kept current every morning