Home › Use Cases › Approved-source versus PO check
Use caseUC0411
🧪 Use-case kit · runnable

Approved-source versus PO check

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A buyer places a purchase order on a supplier, and the approved-source list says whether that supplier may make that part. The list is keyed on a PAIRING — this supplier, this part number — and it carries an approval window, a coarse scope code, a status word, a value ceiling and, beside each entry, a free-text source note. Four of those five are columns and are a lookup. The note is not: a suspension raised since the list was re-issued lives there while the Status column still reads Active, one that was LIFTED lives there while the column still reads Suspended, a written scope that narrows the coarse code lives there as a list of what IS approved, a revision limit lives there and nowhere else because there is no revision column at all, and a distributor's authorisation letter lives there naming the one manufacturer that distributor may supply for. A buyer under schedule pressure reads the columns. Joining every purchase-order line to an approved-source list extract by supplier code AND part number, to the order's own date against two printed dates, to a coarse scope code against the process the line requires, to a per-line value ceiling — and then reading the source note beside each entry for a suspension, a lifting, a written process list, a revision limit and a distributor's letter, by eye, before placement, sixty orders deep.

Audience

A buyer's source-control desk, reading a requisition packet BEFORE the order is placed. The output is what somebody reads before anybody decides anything; the decision is theirs. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual requisition packets, each one purchase order before placement

The corpus is 62 requisition packets, each one purchase order before placement, 0.20 MB (json 3 · jsonl 1 · md 2 · txt 62). Because the shape of an approved-source failure is not the lookup, it is the SOURCE NOTE, and a corpus has to be able to separate the two. 248 of the 277 lines here are settled by a pairing, two date comparisons, a coarse scope code and one multiplication, and free code gets every one of them. The other 29 have an answer that exists only in a sentence, and 23 of those sentences occur exactly ONCE in the whole corpus — which is what makes the mechanical de-memorisation of the note floor the honest comparison rather than the memorising one.

The corpus

  • The 62 requisition packets, each one purchase order before placementgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 62 packets, data/sources.json and the whole answer key are generated from one seed and ship in this repository.

Swap this folder for your own material and the kit is pointed at your requisition packets, each one purchase order before placement. That is the whole change — there is no database to migrate.

One requisition packets, each one purchase order before placement, as the model receives itPO-0001.txt · 1 of 62
REQUISITION PACKET - APPROVED SOURCE CHECK BEFORE PLACEMENT

PACKET HEADER
  Packet             PO-0001
  Purchase order     PO-51000
  Order date         2026-01-12
  Buyer              Halcyon Aerostructures, Kestermill site
  Programme          MERIDIAN-4 propulsion
  Raised by          M. Hallberg, procurement

SOURCE CONTROL PROCEDURE
  Procedure               ASL-2026
  List revision           ASL-2026 list revision 9
  List issued             2025-12-03
  Ceiling basis           per purchase-order line, quantity times unit price
  Scope codes             all-processes, fabrication, special-process, distribution

PART MASTER
   Part         Description                            Approved manufacturer            Revision released
   ANV-10000    Bracket, forward avionics rack         Calderhaven Industrial           D
   PNL-10061    Mount, environmental control unit      Briarfoot Industrial             E
   SPR-10122    Valve body, bleed air                  Kelvinmoor Technologies          F

APPROVED SOURCE LIST EXTRACT
   Code     Supplier                         Type          Part         Scope            Status     Approved from  Expires      Ceiling
   S-2723   Hesketh Industries               manufacturer  SPR-10122    fabrication      Active     2025-01-19     2026-05-28   $5,700.00
   S-3677   Briarfoot Industrial             manufacturer  ANV-10000    fabrication      Active     2025-03-18     2026-04-12   $1,800.00
   S-8228   Calderhaven Industrial           manufacturer  PNL-10061    special-process  Active     2025-02-17     2026-05-05   $3,600.00

PURCHASE ORDER LINES
   #  Part         Rev  Process required             Supplier  Qty    Unit price    Committed      Source note

Abridged — the file continues.

The outcomeWhat a good result looks like

Every purchase-order line carries a verdict, the term of ASL-2026 it rests on, the value at risk to the cent and one row copied verbatim from the packet as evidence, and the order carries CLEAR or HOLD with the held lines named and the total stated. On 62 packets and 277 lines the paid call, with ASL-2026 re-applied in code to its own two readings, takes 235 of 277 line verdicts (84.8 pct) and 25 of 62 orders.

And when it cannot

⚠︎ THE PAID CALL LOSES TO THIS KIT'S OWN FREE NOTE FLOOR AND THAT SENTENCE GOES FIRST. phrase-generic — the note-reading floor with every memorising pattern removed MECHANICALLY, 0 calls, $0.00 — takes 255 of 277 line verdicts against the call's 235, ahead on 21 discordant lines and behind on 41, exact two-sided p = 0.015134. The phrase floor, written with the answer key open, takes 277 of 277 and 62 of 62 orders and the call cannot be told apart from it at all. AND THE OPPOSITE IS TRUE ON THE POPULATION THE CALL EXISTS FOR: on the 29 lines whose answer lives in a source note the call takes 28 against the column floor's 0 and the de-memorised note floor's 7, p < 0.000001 against both. The order-level number is 25 of 62 and that is not a typo: order_all_correct requires every line right on all six graded fields AND the recommendation, the held lines and the total, and the call quoted a row on 68 of the 233 lines the list supports where the answer contract says null. The station does not repair a citation, deliberately.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your approved-source list is clean and your arguments are about pairings, expiry dates and value ceilings. — the free column floor alone — python3 -m evals.run --floor rules
    Every one of those is a lookup, two date comparisons or one multiplication. The floor is 248 of 277 line verdicts and 33 of 62 orders for $0.00 and no network, and it never holds a line the list supports.
  • Your source-control file carries note traffic: suspensions raised between list re-issues, scope narrowed in a sentence, revision limits nobody columnised, distributor letters. — the paid call, and then the free floor over the top of it
    The call is the only arm that reads a note at all: 28 of the 29 lines a note decides, against the column floor's 0, p < 0.000001. Running the free floor afterwards is what takes back the 38 supported lines it holds for nothing.
  • You want to know whether a source-control desk can stop reading notes. — no
    The call missed one of the 11 standing moves and, more to the point, this corpus's notes are drawn from 132 distinct strings. A real file's note traffic is unbounded prose and nothing here measures that.

And where nothing here is good enough:

  • You want one number to route on. — neither arm alone
    The corpus separates the arms on 29 of 277 lines and the other 248 are a majority class. Three of the percentages on this page are unreadable without the sub-population beside them.

At a glanceHow the whole thing runs

85%line verdict correct rechecked pct
2,582 msp50, end to end
$1.20per 1,000 requisition packets · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported requisition packets and data/sources.json at your own part master and approved-source list extracts, keyed on the same packet ids. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need — and paying it on every line of every order. That is the case against the best-fitting scenario (“Your approved-source list is clean and your arguments are about pairings, expiry dates and value ceilings.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose panels are not the eight this parser knows. src/rules.py splits on PACKET HEADER, SOURCE CONTROL PROCEDURE, PART MASTER, APPROVED SOURCE LIST EXTRACT, PURCHASE ORDER LINES, BUYER REMARKS, SIGN-OFF and END OF PACKET, and a missing heading yields an empty section rather than an exception. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from a real difference. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-approved-source. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all four free floors score in-process against the shipped key in a second or two, and python3 -m src.app serves the whole board with the committed run replayed at $0.00.

A living map of modern AI — kept current every morning