Home › Use Cases › Closing an AFE against the capital expenditure procedure the package itself prints
Use caseUC0208
🧪 Use-case kit · runnable

Closing an AFE against the capital expenditure procedure the package itself prints

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An AFE closeout is a reading problem sitting on a spreadsheet problem, and the spreadsheet half is the larger one. Somebody has to go down a cost detail and a commitment register and decide, per line, whether a FINAL vendor document supports it, whether that document was raised against this AFE, whether the line's class agrees with its cost code and whether the charge fell inside the trailing-charge window — then add up what is genuinely this AFE's cost, compare it with the authorisation, and decide whether the overrun forces a supplemental AFE before the well can close. The tie-out everybody already has answers a different question: is there a document identifier in the column. A pro-forma has one. A field ticket has one. An invoice a credit memo reversed three weeks later still has one. The line-by-line closeout tie-out an accountant works against the capital expenditure procedure — reading each vendor document to see what kind it declares itself and which AFE it names, checking the cost-code class, the trailing window and the release documents, then re-adding the actual cost with the charges that are not this AFE's taken back out. Not the decision to close: all four closeout actions are recommendations to the people who take them, and nothing in the kit posts a journal, releases a commitment or raises a supplement.

Audience

The operator's closeout accountant, and the joint-venture team who answer the non-operating partners for what the AFE finally cost. The decision in front of them is not 'is the model accurate'; it is 'which of these findings can I put in front of a field office and a partner without being wrong about the arithmetic underneath it'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual AFE closeout packages

The corpus is 40 AFE closeout packages, 0.28 MB (txt 40). Because the traps are the product. Four of the thirteen patterns are items that LOOK like findings and are not, or the reverse: a charge dated after operations complete that falls inside the 60-day window; a commitment showing an unreleased balance with a filed release document covering it in full; a real finding under the de-minimis threshold on an AFE that is not audit-listed, so it does not block; and the IDENTICAL finding at the IDENTICAL size on an AFE that IS listed, so it does. The last pair are indistinguishable on every printed table and differ only in whether the AFE appears on a list carried on the header for six packages and in the correspondence alone for six more. A checker that holds every package with anything imperfect about it scores well on the nine real patterns and fails all four traps — and that is a checker nobody in a field office keeps reading.

The corpus

  • The 40 AFE closeout packagesunder its source's terms — generated by tools/build_corpus.py from a fixed seed (SEED = 20260830); byte-identical on rebuild, verified by hashing every file across two consecutive builds.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your AFE closeout packages. That is the whole change — there is no database to migrate.

The outcomeWhat a good result looks like

One package is reconciled in about 101 seconds into three answers at once: a disposition per item with the clause the package's own procedure prints for it, a separate blocking call per item, and five reconciliation figures with the closeout action they imply. On the 40-package set the paid arm got 260 of 260 items right, 100 of 100 blocking items caught, all five figures exact on 40 of 40 packages and the closeout action right on 40 of 40 — including the 6 packages sitting on EXACTLY CE-4.1's 10.00 per cent threshold, which the clause says do not need a supplement.

And when it cannot

⚑ THE 100 PER CENT IS THE PROBLEM, NOT THE PRODUCT. One item of 260 was wrong on the model's own answer — AF-0020's CL-0020-04, correctly called RECONCILIATION_INCOMPLETE and then marked BLOCKS_CLOSEOUT, when CE-5.1 says a document the register names and the package does not carry is recorded and requested, not held against the AFE. Pure code caught it and the rechecked column is 260 of 260. That is the whole measured failure surface of this run, which means this corpus has NO HEADROOM LEFT to measure a better model, a worse one, or a cheaper tier against. A kit that scores 100 per cent has stopped being an eval and become a demonstration, and the honest reading is that the next thing to build is a harder corpus, not a better arm. ⚠︎ AND THE RAW 100 PER CENT IS ONE PASS OF TWO. This run's cache records 70 calls against the 40 published: 30 packages were answered twice on identical prompts and the first set was dropped. Scored over those 30, the discarded pass is 98.5 per cent on items and 29 of 30 packages exact, missing 3 evidence cells and AF-0009's actual cost by $414,200.00. The rechecked column is identical on both passes — so the half of this kit that survives a re-fire is the procedure in code, and the reading is the half that moves.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Everything that settles your closeout is printed in a table — procedure-sweep alone — the nine clauses in pure code, no key
    248 of 260 items, 40 of 40 authorised totals, all 10 real supplements caught, for $0.00 and no network.
  • Your reversals, releases and audit listings live in correspondence — The model's reading, with the procedure re-applied in code
    That is exactly the 12 items and 6 blocking calls the floor cannot reach — 0 of 12 on credit reversals and 4 of 10 on audit-listed blocking, against 12 of 12 and 10 of 10.
  • A false supplemental AFE is the expensive failure — Either arm that re-derives actual cost from dispositioned items
    0 false supplements on both model arms; 8 on procedure-sweep and 16 on register-tieout, every one of them a well told it needs partner authorisation it does not need.

And where nothing here is good enough:

  • You want to know whether a cheaper model would do — Neither — this corpus cannot answer that
    The scored arm is at 100 per cent on every denominator on the RECHECKED column, and its raw column scored 100.0 on one pass and 98.5 on the other over the 30 packages fired twice. There is no headroom to separate a better model from a worse one, only one model was run, and what spread the set can show is smaller than the difference any model comparison would need.

At a glanceHow the whole thing runs

99.6–100%the blocking-call agreement over the 260 register items
101,071 msp50, end to end
$42.07per 1,000 AFE closeout packages · Google Gemini 3 Flash

Run once, for real, on 2026-08-30. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/gold.jsonl, then rewrite src/afe.py's regexes and src/segment.py's headings to your layout. Every rate here stops being true the moment the corpus changes, and the floors' rates stop being true fastest — procedure-sweep is the SAME ENGINE the paid arm's recheck uses, so it is measured on a corpus whose printed tables were generated to be internally consistent to the cent. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call to re-derive what a printed column already states. On this corpus that is 95.4 per cent of the work. That is the case against the best-fitting scenario (“Everything that settles your closeout is printed in a table”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A package that is not laid out like these. src/afe.py is a set of regexes written for these underlined headings, these fixed-width tables and these identifier shapes; a row it cannot recognise is SKIPPED rather than guessed at, because a mis-parsed cost line silently changes both a denominator and a total. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?One model, one key, one day, one SCORED run. Provider-side reasoning is left at its default and re-rolls per call, so a second run of the identical prompts does not necessarily produce identical replies. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-30 — r001-afe-closeout. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured builds the corpus, grades the answer key through the arms' own parser, runs BOTH free floors over all 40 packages, and serves the whole app with the recorded run replayed inside it — every number on this page except the two model columns and the injection probe. requirements.txt names nothing, so the fork test is git clone and python3. The only control that needs a key is one button, and with no key it returns 200 saying so rather than failing.

A living map of modern AI — kept current every morning