The business caseThe problem this solves
A customs information query arrives as a letter about one entry. It is prose: ten or eleven sentences of which some carry a request, some carry several, some are courtesy and some are background that a request is buried behind. A broker reads it, works out what is actually being asked and at what scope — the whole entry, or one invoice line — then goes through the entry file for the record that answers each one, and drafts a reply. Where the file has nothing, the reply cannot be sent at all: the request has to be named and the packet held. The first pass a broker makes over a customs information query — reading the request set out of the letter, matching each request to a record on the entry file, and drafting the reply or the hold. A licensed broker still decides what is sent.
Audience
The import desk that has to answer these letters, and whoever decides whether a drafting step is worth buying. The answer this page gives is a qualified yes: it beats the free floor a desk could build, and it ties a free ceiling that had read the corpus's own generator. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual query packets
The corpus is 64 query packets, 0.26 MB (json 5 · jsonl 1 · md 2 · txt 64). Because the reading this kit pays for only exists in PROSE. 573 sentences carry 437 requests over 64 packets: 23 sentences carry more than one request, 8 requests span two sentences, 79 courtesy sentences ask for nothing and 16 background sentences sit in front of a buried request. Eight letter shapes by eight file profiles, EACH COMBINATION EXACTLY ONCE, so no letter shape is a status's fingerprint — tools/build_corpus.py --check-leak asserts that over 7 groups. And the filing system's own suggestion is wrong on 40 of the 64 packets, which is what makes reading the file worth anything.
The corpus
- The 64 query packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your query packets. That is the whole change — there is no database to migrate.
CUSTOMS QUERY PACKET
PACKET HEADER
Packet CQP-0001
Importer Torrendale Housewares, importer code IM-2088
Supplier Ondreval Manufacturing
Entry ENT-2026-110004
Query reference CIQ-40011
Query received 2026-07-15
Purchase order PO-1003, placed 21 April 2026
Transport document BL-1005, booking BK-2200
Broker Halverd & Coine import desk
Send gate A licensed customs broker decides what is sent. This packet
drafts a reply and transmits, files, releases and clears nothing.
DECLARED BROKERAGE STANDARD
Response standard HCB-RS-2026
Revision 2 in force from 2026-06-01
Reply body limit 700 characters
These values are the brokerage's own and are supplied by the operator for this
scenario. They are not claimed to be any authority's requirement, and no authority's
rule, threshold or period is stated anywhere in this packet.
CUSTOMS INFORMATION QUERY — as received, one sentence per line
Issued to the broker of record for the entry named above.
[s1] This query concerns entry ENT-2026-110004, filed by your brokerage on behalf of Torrendale Housewares.
[s2] Please quote the reference above on every page of your reply.
[s3] 1. Please provide the written agreement between the parties covering the price of item 41205 on line 3.
[s4] 2. Please provide the insurance record for the shipment carrying this entry.
[s5] 3. Please provide the packing list showing the packing detail for item 71402 on line 2.
[s6] 4. Please provide the freight invoice for the shipment carrying the goods entered under this entry.
[s7] 5. Please provide the manufacturer's production records for line 2.Abridged — the file continues.
The outcomeWhat a good result looks like
Every request the letter makes is on the draft, each one answered from a record printed on this entry file, and any request the file cannot answer is named and the packet held rather than sent.
And when it cannot
A request read out of a courtesy sentence, two requests in one sentence merged into one, or a reply sent that should have been held. Measured: the request set is wrong on 3 of 64 packets and the route on 8 of 64; the status is wrong on 0.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- The query letters are prose and one sentence can carry several requests — the paid call, then the station
this is the only thing it buys and it buys it decisively — the request set is exactly right on 61 of 64 packets against the domain floor's 7, and on the 40 packets the filing system's suggestion gets wrong, 39 against 3 (both p < 0.000001). - The letters come off a small number of fixed templates you control — free code — write the patterns
that is exactly whattunedis, and it TIES the paid call on all five paired claims (request set 64 against 61, p 0.250000) for $0.00. - You only need to know whether a packet can be answered at all — the station, free, over any arm's request set
over the null arm's own request setsrc/recheck.pyalone recovers 18 of 24 unanswerable requests for $0.00, by looking up whether a cited record is printed on the packet.
And where nothing here is good enough:
- You want a headline that says the model is accurate — none of the three obvious ones
disposition_accis a CLASS PRIOR — 365 of 437 requests are RESPONDS, so the accept-everything constant scores 84.1 pct against the floor of record's 75.3.citation_accandpair_agreementare ~100 pct for every arm. All three arenever_a_headlinein data/floors.json and in the run record.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own query packets and data/gold.jsonl with one JSON line per packet carrying the request set, each request's disposition and answering record, the status, the held requests and the route. Every measured figure on this page stops being true the moment the corpus changes. Corpus lens → |
| When is this the wrong choice? | Avoid: Reading the improvement as an improvement in the DRAFT. Status is right on 62 of 64 for the free floor too; the gap is the reading, not the letter. That is the case against the best-fitting scenario (“The query letters are prose and one sentence can carry several requests”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A letter materially wider than this corpus's widest. 11 requests over 10 sentences is the top of the range here, and it is the packet the request set gets wrong (2 read that are not requests, 2 missed) and the one reply whose body went over budget. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THE PAID CALL DOES NOT BEAT tuned ON ANY CLAIM, AND THIS SET CANNOT TELL THEM APART. 61 against 64 on the request set, p 0.250000; 0 of 5 paired claims significant. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-21 — r001-customs-query. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board and scores every free arm offline: python3 -m evals.baseline computes all seven arms at $0.00 with no network, python3 -m evals.run --stub exercises the whole harness end to end without a socket, and python3 -m src.app serves both surfaces — the draft button is disabled and says why. The 13 screenshots in docs/shots/ were taken against exactly that no-key board. What a clean checkout cannot do is reproduce the paid column: that needs a key, and re-running it is a new run id.












