The business caseThe problem this solves
A shopper typed something into the storefront's search box and got nothing they could buy - no results at all, or results nobody wanted. The session log says what happened; it does not say why. Six different things look identical from the outside: the shopper was not asking for a product, the category genuinely carries no matching item, the item exists but has no usable index record, the shopper's word is not a word the index knows, the item is there but lacks the attribute they filtered on, or a ranking rule buried it. Each one is a different team's job. Telling them apart means opening the search configuration, the live index records, the FULL catalog extract for the category and its completeness certificate, the synonym and ranking rules, both the merchandiser's and the analyst's notes and the query's prior occurrences - eight panels - and reading them in the order the desk's card says. the analyst's read of eight panels per failing query to decide which team owns it - it publishes no synonym, writes no attribute, rebuilds no index, changes no ranking rule and rates no merchandiser.
Audience
A retail search-operations lead deciding whether to put a model in front of a queue of failing queries. This report's answer is NOT ON ACCURACY: on the cause the paid call does not beat the free walk this kit ships, it loses the recurrence column and the complete row to it outright, and the one slice that looks like a win is beaten by a constant. What is worth taking is the framework, the cost of running it, and the four packs no free code reaches. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual failing onsite search queries
The corpus is 64 failing onsite search queries, 0.20 MB (json 5 · jsonl 1 · md 1 · txt 64). A real null-result triage pack carries a retailer's live search configuration, its full catalog extract and its query log. All three are commercial documents, the query log is shopper behaviour, and a redacted copy cannot be labelled without the labels leaking what was redacted. Each file is built as a structure and the answer key is SRT-2026 applied to that same structure, so every cause is derivable rather than somebody's opinion.
The corpus
- The 64 failing onsite search queriesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your failing onsite search queries. That is the whole change — there is no database to migrate.
ONSITE SEARCH FAILURE SNT-0001
====================================================================================================
PANEL 1 - THE QUERY AND THE SESSION
Retailer Harrowmere Outdoor & Home (synthetic retailer)
Storefront SF-US (synthetic storefront)
Category CAT-11 rain and shell jackets
Device class desktop
Query date 2026-05-04
Query as typed rain shell
Concept asked for rain shell - the intake's normalisation of the query
Filters applied none
Results shown 0
Session ended refined the query and left
PANEL 2 - WHAT THE SEARCH ENGINE RETURNED (first page, in rank order)
no results on the first page
PANEL 3 - SEARCH INDEX RECORDS FOR THIS CATEGORY
SKU-40002 status indexed 29 April 2026 from build CB-100
SKU-40002 indexed text lightweight rain jacket
SKU-40002 attributes colour=slate returnable=yes
SKU-40002 stock state in stock
SKU-40003 status present in build CB-102, built 27 April 2026
SKU-40003 indexed text wind shell
SKU-40003 attributes colour=slate returnable=yes
SKU-40003 stock state low stock
PANEL 4 - FULL CATALOG EXTRACT FOR CAT-11 (every item on the category, sellable or not)
not attached to this pack
PANEL 5 - SEARCH CONFIGURATION IN FORCE (the rule table as the profile prints it)
Rule Kind As printed in the profile
SR-10 synonym maps "anorak" to "rain shell"
BR-05 ranking buries items in stock state "low stock" below the first page
SR-50 synonym maps "plimsoll" to "trail runner"
PANEL 6 - MERCHANDISER NOTES ON THE CONFIGURATION (free text, as typed)
SR-50 has not been published to the live index.
Abridged — the file continues.
The outcomeWhat a good result looks like
One row per failing query an analyst could work as written: the cause under SRT-2026, the printed line of the pack that establishes it, the recurrence verdict against the operator's threshold, and - derived in pure code, never asked - the owning queue, the fix kind and the recurrence count.
And when it cannot
It abstains. A pack whose full catalog extract never came through cannot be walked, and the answer is NEEDS-REVIEW whatever either note says - 3 of 64, and every arm gets all 3 because the rule is structural in src/recheck.py. Where the operator supplied no recurrence threshold for a category it answers the literal threshold not supplied and never borrows one - 16 of 16 with 0 breaches, a number a constant emitting that literal would also reach, so it is reported and never carded alone.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- your merchandiser and analyst notes are written in the card's own words, and your export prints the catalog extract and its certificate — the free reading floor - evals/baseline.py, $0.00
it takes 8 of the 8 packs a standard-worded sentence decides and 3 of the 3 the printed panels decide alone; the paid call takes 6 and 3. A call there buys nothing and loses packs a written vocabulary already had. - your merchandisers word their certificates, index notes and rule scopes in their own words — STILL measure before you buy - and measure against a constant as well as your vocabulary floor
this is the slice that looks like the money's job and on this corpus it is not: the paid call takes 6 of the 18 paraphrase-decided packs against the floor of record's 2, but a free constant that reads nothing takes 10 (p = 0.2188). A 6-against-2 quotation would have been an artifact of which arm was called the floor. - the recurrence verdict matters, because a pattern triggers a category review — free code, unconditionally
SRT-9 over PANEL 8 is a count against a printed threshold. Three separate free arms score it 64 of 64 and the paid call scores 52 (p = 0.0005 against it). There is no version of this column worth paying for. - you need the four packs no free arm reaches - a paraphrased index status, and one paraphrased vocabulary pack — the paid call on those packs only, and only after your own floor has settled the rest
SNT-0021, SNT-0024, SNT-0027 and SNT-0055 are the whole measurable value of the call on this corpus: 4 of the 9 packs beyond every free arm on the cause. Three of the four are the paraphrased index statuses phase 1 named as hard.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/policy.json with your own triage card - the ordered tests, what your completeness certificate licenses and what it does not, the evidence line per cause and the recurrence thresholds your operator actually supplies - and src/taxonomy.py's causes and their ORDER with yours. The measured result does not travel. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per query for a date comparison and a word list. That is the case against the best-fitting scenario (“your merchandiser and analyst notes are written in the card's own words, and your export prints the catalog extract and its certificate”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a merchandiser or analyst note worded in a way nobody anticipated. The free floor of record reads the usual words for a certificate that stands, a stale index record, a live synonym and a scoped ranking rule and takes 2 of the 18 paraphrase-decided packs; the paid call takes 6 - and a constant that reads nothing takes 10. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | The recurrence column is FREE CODE at 64 of 64 and may never be quoted as a kit result. Three separate free arms score it perfectly - the domain floor of record, the tables-only walk and the note-follower - so headroom above free code there is 0 packs, and the paid call scores 52 (p = 0.0005 against it). 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak tariff - THE RUN OF RECORD. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r002-search-null-triage. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — The kits repository is private, so this is a record of what was run rather than an offer: on a copy of this folder with no key configured and no .env above it, all four free floors (45, 27, 15 and 10 of 64 rechecked), evals/check_labels.py (0 disagreements), python3 -m evals.paired --verify (221 derivations, 0 mismatched) and a re-score of both committed paid runs from their own caches reproduced at 0 calls and $0.00, and the board on port 9488 served every route. Re-running a paid arm needs a provider key. ⚠︎ THE REPLY CACHES SHIP, AND THAT WAS MEASURED RATHER THAN ASSUMED: with results/cache-*.jsonl removed into a scratch copy the board's peak-tariff row renders as an em dash, because src/app._reprice() returns None without them - so the caches are evidence a shipped surface reads and the kit's .gitignore line that would have dropped them was removed.









