Home › Use Cases › Read a customer's own documents into requested-specification rows
Use caseUC0359
🧪 Use-case kit · runnable

Read a customer's own documents into requested-specification rows

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A customer's specification arrives as five documents that disagree about what counts. The requirement a chemist has to write down may be a labelled line on the RFQ, a sentence in the body of a packing note, a number the customer says is typical rather than required, or the same limit twice with two different values - and the one that is read past is the one nobody ever tests to. Re-reading five documents by hand to build one specification, and the second read-through that finds the line the first one missed.

Audience

The customer-service desk that opens a specification, and the quality manager who accepts it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual specification packets

The corpus is 62 specification packets, 0.12 MB (txt 62). A customer specification names the customer, the product, the grade and the limits they buy against; it is the most commercially sensitive document either side holds and one of the two parties did not choose to be in a dataset. No such corpus can be published, so this one is manufactured from a fixed seed and declared as manufactured. What it exercises is a READING problem with the parsing problem deliberately removed: a fixed layout, so the comparison between a phrase list and a reading is about the reading.

The corpus

  • The 62 specification packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your specification packets. That is the whole change — there is no database to migrate.

One specification packet, as the model receives itSPC-0001.txt · 1 of 62
CUSTOMER SPECIFICATION PACKET   SPC-0001
==============================================================================

ENQUIRY FACTS
  Enquiry reference     ENQ-2026-1107
  Customer              Brightwater Coatings Ltd
  Delivery site         Ashgrove Works
  Product               Adhanol 400
  Grade requested       technical
  Annual volume         40 tonnes
  Enquiry received      2026-02-02
  Requirements noted on the request  See the documents logged against this enquiry.
  Receiving desk        Customer service

DOCUMENT INDEX -- what arrived with the enquiry, and the type it was logged under
  D1   RFQ cover sheet                      received 2026-02-02
  D2   Emailed specification sheet          received 2026-02-02

DOCUMENT D1 -- RFQ cover sheet
  Issued by      Brightwater Coatings Ltd, technical purchasing
  Covers         enquiry ENQ-2026-1107
  Reference      ENQ-2026-1107-D1
  Product enquired for                                  Adhanol 400
  Grade                                                 technical
  Certificate required                                  a certificate must accompany each consignment

DOCUMENT D2 -- Emailed specification sheet
  Issued by      Brightwater Coatings Ltd, quality department
  Covers         enquiry ENQ-2026-1107
  Reference      ENQ-2026-1107-D2
  Sheet issued                                          2026-02-03
  Applies to                                            Adhanol 400
  pH, range                                             5.0 to 6.5 at 20 pct in water by TM-2500

INTAKE NOTES
  Logged to the enquiry: "Documents received and filed against the enquiry reference."

The outcomeWhat a good result looks like

The requested-specification rows a chemist can act on, with the line each was read from - plus what could not be resolved and what was stated twice with different numbers, named rather than silently picked.

And when it cannot

A specification opened without a limit the customer asked for, or opened carrying one of two numbers with nothing to say the other exists.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your packets are fixed-layout and your customers write limits the way your procedure does — the free rules floor
    it scores 46 of 62 whole packets here against the paid call's 39, for $0.00 and no key, and it takes 9 of 10 tabular packets and 8 of 9 unresolved ones
  • Your customers attach last year's email thread, or say a family does NOT apply to this grade — the model call
    the floor takes 0 of the 6 historical packets and 5 of the 7 negated ones; the call takes 3 and 7. Across both it captures a requirement nobody asked for on 3 packets against the floor's 14.
  • What you most need is that a requirement stated twice with two different numbers is never quietly resolved — the free rules floor, and a human step behind whichever arm you run
    the floor takes 62 of 62 conflict sets exactly right and the call 57. CSI-3 cannot help: it needs the conflict to have been CAPTURED, so one nobody read never fires it and the specification goes to the quality manager carrying one of two numbers.
  • What you most need is not to email customers about limits they already sent — either arm, and read the unresolved column rather than the headline
    the call chases 5 of the 30 clean packets and the floor 6. It is close, and it is the only failure here that generates outbound contact.
  • Your requirements arrive as sentences in documents nobody indexes - a packing note, a certificate template, a drawing sheet — either arm, and NOT a capability-table filter
    both real arms read every document body and take 10 of 10 prose packets; the null floor, which is what a filter that trusts the desk's index degenerates to, misses all 10. Under a note asserting those documents carry nothing, the call still misses 0 of 10.

At a glanceHow the whole thing runs

63%packet all correct pct
1,733 msp50, end to end
$1.39per 1,000 specification packets · openai/gpt-5-6-luna

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packets in the same five-panel shape and data/enquiries.json with your own register rows, then re-run both free floors for $0.00 before spending anything. The answer key does not come with them. Corpus lens →
When is this the wrong choice?Avoid: A model call, until you have measured the floor on YOUR packets. It is the arm this kit could not beat. That is the case against the best-fitting scenario (“Your packets are fixed-layout and your customers write limits the way your procedure does”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A specification that arrives as a spreadsheet attachment or a scanned drawing - this kit reads text and there is no OCR anywhere in it. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?ONE scored run, one model, one tier. Nothing here says whether a larger model asks the section 3 actionability question better - and that is one of the two fields this arm loses on, so it is the commercial question and it is open. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-spec-intake. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board - all 62 packets, both free floors, CSI-2026, the capability index, the citation scorer and the committed run replayed - and scores both floors offline. The model button is disabled and says why. Measured on this machine with API_KEY blanked, which is also how every screenshot was taken.

A living map of modern AI — kept current every morning