Home › Use Cases › Restriction and length-of-stay audit
Use caseUC0378
🧪 Use-case kit · runnable

Restriction and length-of-stay audit

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A revenue manager writes the season's strategy in prose — hold a two-night minimum over the festival weekend, open arrivals once the group block releases, stop-sell everything but the suites — and somebody then loads it into the booking system as a grid of dates, room types and restriction values. The two drift apart immediately: a correction is typed into the memo and never loaded, a stop-sell from last season is never lifted, a minimum stay is loaded a day wide. Nobody reads the memo back against the grid, because the memo is paragraphs and the grid is hundreds of cells. Reading a season's strategy note back against the restriction grid by hand, date by date and room type by room type, which nobody does and which is why the drift is found by a guest failing to book.

Audience

A revenue manager or commercial director deciding whether a model is worth buying for this job, and an engineer deciding what to build. The honest answer this kit reaches is a qualified no on the headline and a clear yes on one specific half of the work, and the board opens on the half that does not flatter it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual property packets

The corpus is 40 property packets, 0.06 MB (json 3 · jsonl 1 · md 2 · txt 40). Because the difficulty had to be REAL and had to be MEASURABLE, and real hotel memos are neither public nor licensable. Each packet carries at least one prose construct a date parser can reach and at least one it cannot, and data/corpus-stats.json counts how many graded rows each construct accounts for — so the board can publish accuracy BY CONSTRUCT rather than one number averaged over two populations that behave nothing like each other. The first cut of this corpus was fully solvable by the free floor (307 of 307) and was rebuilt; that is recorded here rather than quietly fixed, because a corpus a rule reader clears completely measures the author's parser and not the model.

The corpus

  • The 40 property packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your property packets. That is the whole change — there is no database to migrate.

One property packet, as the model receives itRA-0001.txt · 1 of 40
PROPERTY: Harbour Crest Hotel (HCR)        PACKET: RA-0001
HORIZON: 2026-05-04 to 2026-05-15
ROOM TYPES: STD (Standard King); DLX (Deluxe Harbour); SUI (Harbour Suite); FAM (Family Room)

==============================================================================
PART A - REVENUE STRATEGY NOTE, as written by the revenue manager
==============================================================================

  The Ashcombe group block releases on 12 May.

  A1. If the Ashcombe block has not picked up by 7 May we hold a three-night minimum from that date to the end of the horizon. It has not picked up.

  A2. Hold a two-night minimum from 11 May to 12 May on the Harbour Suites only.

  A3. Open arrivals again from the release date to the end of the horizon — we need the shoulder nights back.

  A4. Drop the seven-day advance fence from 8 May onward; we would rather take the booking than protect the rate.

  A5. Apply the seven-day advance fence over the same dates as A2.

  A6. Last year we ran a four-night minimum over this period and it cost us eleven room nights; we are not repeating it.

  A7. Watch the OTA parity report daily — last quarter we were undercut on two channels and nobody noticed for nine days.

==============================================================================
PART B - RESTRICTIONS AS LOADED IN THE BOOKING SYSTEM
==============================================================================

   #  DATE        ROOM  RESTRICTION  LOADED
   1  2026-05-11  FAM   STOP_SELL    on sale
   2  2026-05-11  STD   CTA          open
   3  2026-05-12  DLX   CTA          open
   4  2026-05-12  DLX   FENCE        none
   5  2026-05-12  STD   MINLOS       3 nights
   6  2026-05-15  FAM   FENCE        none

The outcomeWhat a good result looks like

Every date where the loaded configuration contradicts the stated strategy, named, with both sides quoted — plus the restrictions that are loaded and that nobody asked for, which is the half a human audit misses because there is no sentence in the memo to find them by.

And when it cannot

It attaches a directive to a date that directive never reaches, and then rules correctly on the wrong premise. 30 of 322 rows came back as a finding invented — a ticket a revenue manager has to open and close — and 22 as a real finding missed, which is a night the property keeps selling wrong.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your strategy notes are written as explicit date ranges — the free resolver floor — evals/baseline.py, 0 calls, $0.00
    it is PERFECT on explicit ranges, open-ended dates, named periods, corrections, room carve-outs and days of the week — 10 constructs, no misses — and it is a file you can read in twenty minutes
  • your notes reference other directives, resolve conditions later, or describe periods relative to an event — the paid call
    +44 rows over free code on those 80 rows, and the floor scores 7 of 80 on them — these are the constructs a date parser cannot be written for
  • you want the best answer available and cost is not the constraint — route in code — the floor on rows it resolves, the call on the rest
    the two arms read the same NUMBER of rows correctly on different rows, which is the textbook condition for routing to beat either
  • you need the audit to be checkable by a person afterwards — either arm, with src/policy.py doing the ruling
    every published cell is recomputed from two readings by pure code and every finding carries the sentence it rests on

At a glanceHow the whole thing runs

71%governs correct pct
2,020 msp50, end to end
$0.75per 1,000 property packets · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packets in the same two-part shape — the note as written, then the loaded table with a #, a date, a room code, a restriction kind and the loaded value — and write data/gold.jsonl with the governing directive and the requirement for each row. THE MEASURED MARGIN DOES NOT TRAVEL. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call that scored WORSE than it did on every one of those constructs. That is the case against the best-fitting scenario (“your strategy notes are written as explicit date ranges”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A strategy note whose periods are never defined anywhere — 'the usual festival dates' with no calendar in the document. Every arm returns NONE and every restriction on those dates reads as UNSUPPORTED, which is a page of false findings rather than an error. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the margin survives on REAL revenue strategy notes. The corpus is synthetic because real hotel memos are neither public nor licensable, and the whole comparison turns on the mix of prose constructs — 80 of 322 rows here are governed by a construct free code cannot compose. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN), one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-restriction-audit. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the entire board — the corpus, the answer key, the committed run, all three free floors, the by-construct table, the money and the injection probe — and scores all three free floors offline in pure Python. tools/build_corpus.py --verify reproduced all 40 packets byte-for-byte under a different PYTHONHASHSEED, and evals/check_labels.py passed 2,696 assertions with no network.

A living map of modern AI — kept current every morning