Home › Use Cases › Turn a shelf-bay photograph into a ranked list of planogram exceptions
Use caseUC0432
🧪 Use-case kit · runnable

Turn a shelf-bay photograph into a ranked list of planogram exceptions

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A merchandising auditor walks a store with a photograph of each shelf bay and that bay's planogram -- the plan of which item sits in which slot, how many facings it owns and what its shelf tag should say. They compare the two by eye, position by position: an empty slot where a product should be, a product sitting in another product's slot, a tag price that disagrees with the planogram. Then they decide which of those is worth a visit first. The by-eye comparison of one shelf-bay photograph against its planogram, position by position, that decides which missing facings, misplaced items and wrong tag prices an auditor goes to confirm first.

Audience

A retail operations or merchandising product manager deciding whether an OCR station plus a model belongs in front of shelf-audit photographs. The honest answer on this corpus is no, on every path measured: on perfect text the free regex floor beats the model (exception F1 0.9615 against 0.7941, McNemar exact p 0.000019); Mistral OCR 4.1 returned no shelf text on 56 of 60 photographs; and AWS Textract Detect Text read the words but not one product name whole, so the model on its text scores 0.8303 of positions -- below the 0.8985 of reporting every position as planned -- with exception F1 0.2195. No track clears the 0.9685 accuracy floor. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual shelf-bay photographs

The corpus is 60 shelf-bay photographs, 36.13 MB (png 60 · txt 60). A shelf bay is a place where the thing an auditor judges is spatial -- which slot, how many facings, which tag -- and the only evidence a text station can pass on is words. That makes it a hard, honest test of arrangement A1: every exception is legible as text by construction (a product name printed on each facing, an item code and price on each tag, a bay header), yet the audit depends on alignment the text does not mark -- no empty-space marker, no frame edge. The degradations were planned from the row's own open question, poorly lit and partial-shelf photographs, and every exception code appears on at least 3 bays in every condition.

The corpus

  • The 60 shelf-bay photographsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your shelf-bay photographs. That is the whole change — there is no database to migrate.

One shelf-bay photograph, as the model receives itSB-0001.txt · 1 of 60
AISLE 23   BAY 04   SALTY SNACKS
PLANOGRAM PG-8127  REV A  EFF 2026-08-05
CALLOWEN CHEESE PUFFS JALAPENO 7 OZ
CALLOWEN CHEESE PUFFS JALAPENO 7 OZ
CALLOWEN POPCORN BBQ 13 OZ
CALLOWEN POPCORN SEA SALT 13 OZ
CALLOWEN POPCORN SEA SALT 13 OZ
CALLOWEN POPCORN SEA SALT 13 OZ
CALLOWEN POTATO CHIPS BBQ 9 OZ
CALLOWEN POTATO CHIPS BBQ 9 OZ
CALLOWEN POTATO CHIPS BBQ 9 OZ
#40101 $0.613/OZ $4.29
#40115 $0.292/OZ $3.79
#40122 $0.376/OZ $4.89
#40129 $0.599/OZ $5.39
CALLOWEN PRETZEL TWISTS SEA SALT 7 OZ
CALLOWEN PRETZEL TWISTS SOUR CREAM 7 OZ
CALLOWEN PRETZEL TWISTS SOUR CREAM 7 OZ
CALLOWEN TORTILLA CHIPS BBQ 13 OZ
CALLOWEN TORTILLA CHIPS BBQ 13 OZ
RUSTMERE POPCORN SEA SALT 7 OZ
RUSTMERE POPCORN SEA SALT 7 OZ
RUSTMERE POTATO CHIPS BBQ 9 OZ
RUSTMERE POTATO CHIPS BBQ 9 OZ
#40143 $0.641/OZ $4.49
#40150 $0.670/OZ $4.69
#40157 $0.315/OZ $4.09
#40199 $0.441/OZ $3.09
#40206 $0.588/OZ $5.29
RUSTMERE POTATO CHIPS JALAPENO 13 OZ
RUSTMERE TORTILLA CHIPS BBQ 13 OZ
RUSTMERE TORTILLA CHIPS SEA SALT 9 OZ
RUSTMERE TORTILLA CHIPS SEA SALT 9 OZ
RUSTMERE TORTILLA CHIPS SOUR CREAM 13 OZ
RUSTMERE TORTILLA CHIPS SOUR CREAM 13 OZ
RUSTMERE TORTILLA CHIPS SOUR CREAM 13 OZ
WENDRAKE CHEESE PUFFS SOUR CREAM 9 OZ
WENDRAKE CHEESE PUFFS SOUR CREAM 9 OZ
#40213 $0.353/OZ $4.59
#40220 $0.238/OZ $3.09
#40227 $0.321/OZ $2.89
#40234 $0.330/OZ $4.29
#40241 $0.332/OZ $2.99
WENDRAKE PITA CRISPS SEA SALT 7 OZ
WENDRAKE PITA CRISPS SEA SALT 7 OZ
WENDRAKE PITA CRISPS SOUR CREAM 9 OZ
WENDRAKE PITA CRISPS SOUR CREAM 9 OZ
WENDRAKE PRETZEL TWISTS JALAPENO 9 OZ
WENDRAKE TORTILLA CHIPS JALAPENO 13 OZ
WENDRAKE TORTILLA CHIPS SOUR CREAM 7 OZ
WENDRAKE TORTILLA CHIPS SOUR CREAM 7 OZ
WENDRAKE TORTILLA CHIPS SOUR CREAM 7 OZ
#40255 $0.727/OZ $5.09
#40262 $0.388/OZ $3.49
#40283 $0.554/OZ $4.99
#40297 $0.407/OZ $5.29
#40304 $0.570/OZ $3.99

The outcomeWhat a good result looks like

A ranked list of proposed exceptions for one bay, each one a planogram position and a code -- MISSING_FACING (a void ranks first), PRICE_TAG_MISMATCH, WRONG_PLACEMENT -- derived by src/rules.py from the model's per-position report. In this kit's own units: 60 bays, 1,143 planogram positions, 78 true exceptions (24 missing facings of which 12 are voids, 30 wrong placements, 24 price-tag mismatches), 12 bays with none. The list is advisory; nothing is sent to a store.

And when it cannot

A different failure at each stage. With the perfect reading the model MISSES rather than invents: 54 of 78 exceptions found, 4 false; 20 of its 24 misses are wrong placements, every one half of a two-position swap on one shelf, and 4 are voids. With a real reading station the list goes wrong two ways. Mistral OCR 4.1 returned the bay header and image placeholders on 56 of 60 photographs, the rules put those positions out of frame, and 74 of 78 exceptions never reach an auditor -- the kit never fills that silence with a guess. AWS Textract Detect Text returned the words and no name whole, and the list fills with noise instead: 141 false exceptions (71 missing facings that are not missing, 33 of them on the 20 clean bays; 57 price-tag mismatches; 13 wrong placements) beside 27 of 78 found -- 22 of 24 price mismatches, 0 of 30 wrong placements, 2 of 12 voids.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Planogram audit where a reading station returns clean shelf text in order — src/floor.py -- the regex floor over the station's text, no model call
    on perfect text it finds 75 of 78 exceptions against the model's 54 (p 0.000019) at $0.00 of model spend
  • The same audit on noisy OCR text where false gaps cost an auditor a walk — the model call after the reading, with the rules kept in pure code
    on AWS Textract text the model raised 71 false missing facings against the floor's 1,064 and won positions 914 to 23 (p < 0.000001); on the 4 bays Mistral did read, 0 against 34 and 40 to 1

And where nothing here is good enough:

  • Shelf photographs read by Mistral OCR 4.1 or its batch tier — neither -- change the reading station first
    56 of 60 photographs returned no shelf text on both tiers and both documented image settings; positions 0.0910 and 0.0866, below the do-nothing floor's 0.8985
  • Shelf photographs read by AWS Textract Detect Text — neither arm -- the text does not carry where a word sits
    0 of 1,143 product names arrive whole; the fast tier scores 0.8303 of positions, below the do-nothing floor's 0.8985, exception F1 0.2195, with 71 false missing facings, 33 on clean bays

At a glanceHow the whole thing runs

9–9%field exact match
2,571 msp50, end to end
$2.81per 1,000 shelf-bay photographs · GPT-5.6 Luna

Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Put your photographs in data/corpus/ and write one SB-####-planogram.json per bay to data/planograms/ in the same shape (position ids, item codes, names, facings, price, weekly units), then a data/gold.jsonl line per bay carrying image, planogram and the expected exceptions and verdicts. The answer key does not come with you, and neither do these numbers. Corpus lens →
When is this the wrong choice?Avoid: Do not carry it to tags it was not written for: its tag pattern is this corpus's #<code> $<price> template, and one misread letter in a name turns a facing into a false wrong placement plus a false gap. That is the case against the best-fitting scenario (“Planogram audit where a reading station returns clean shelf text in order”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A photograph the reading station treats as a picture rather than a page. On 56 of 60 bays Mistral OCR 4.1 returned the header and '![img-0.jpeg](img-0.jpeg)' where the shelf is; on SB-0024, a clean full bay, it returned 12 words and no product name, tag code or price, and the rules put 20 fully framed positions out of frame. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the model or the floor finds more exceptions on real OCR text. On AWS Textract Detect Text text the two cannot be separated (17 vs 19 discordant, McNemar exact p 0.867939); on Mistral OCR 4.1 text only 4 bays carried shelf text. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-12 — r003-shelf-compliance-audit-textract. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Observed on this checkout with the machine's own python3, no virtualenv, no install and no credential configured: python3 -m evals.check_labels is CLEAN, python3 -m src.rules --self-test and python3 -m src.measured --self-test pass, python3 -m evals.run re-scores every committed arm and writes nothing it was not asked to, and python3 -m src.app serves the whole board from disk with no key. Rebuilding the photographs needs Pillow (requirements.txt); nothing else does. Any live model or station call needs a key and was not part of this observation.

A living map of modern AI — kept current every morning