The business caseThe problem this solves
A hotel group's distribution desk has to agree every channel booking with the property's PMS: the channel manager says a booking is LIVE for these nights in this room code, and the PMS either holds a reservation for it or it does not. The printed columns decide most of the answer — the booking row, the reservation row, the room type mapping. What decides the rest is a desk note: a cancellation made at the channel after the last message came through, a guest re-dating at the extranet, a mapping change effective for arrivals on or after a date, a booking moved to the sister property by agreement, a channel re-sending a booking it already delivered — often written as 'yesterday' or 'Monday', and often stamped BEFORE the row's last message, in which case it changes nothing. The channel manager's own exception report disagrees with the procedure on 21 of the 64 extracts in this corpus. Opening one property's channel extract, matching every channel booking to the PMS reservation that holds it, comparing state, nights and room type through the mapping table, finding the double holds and the holds on cancelled bookings, and reading every desk note for a change made at the channel after the last message, a mapping change, an agreed move or a re-send.
Audience
The distribution or reservations desk at a hotel group reconciling each property's channel bookings against its PMS before arrivals — and the revenue and front office desks a walk-risk booking is handed to. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual channel extracts
The corpus is 64 channel extracts, 0.29 MB (txt 64). It is generated because it has to be. A real channel extract is a hotel's commercial record: guest names on every booking, rates and commissions, channel contracts, and desk notes written by named people about named guests. None of that can be published, and a published corpus would have had the one thing this kit measures — the sentence that decides a reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 channel extractsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every property, channel, booking, reservation and message is an invented code, every note speaks for a desk, and THERE IS NO GUEST NAME IN THIS CORPUS AT ALL. evals/check_labels.py sweeps all 64 files for a person-shaped name on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your channel extracts. That is the whole change — there is no database to migrate.
================================================================================
CHANNEL BOOKING RECONCILIATION -- ONE PROPERTY, ONE CHANNEL-MANAGER EXTRACT
================================================================================
FILE BKM-0001
PROPERTY HTL-201 (180 rooms)
EXTRACT RUN 2026-07-01 06:00
ARRIVALS WINDOW 2026-07-01 to 2026-07-14
PACK COMPILED 2026-07-01 10:30 (Wednesday)
CURRENCY USD
-- CHANNELS CONNECTED (channel manager) ----------------------------------------
CHANNEL FEED
CH-A real-time
CH-B real-time
CH-C real-time
-- ROOM TYPE MAPPING (channel manager, as configured) --------------------------
CHANNEL CODE PMS TYPE
CH-A ACC-KING ADK
CH-A KING-STD KNG
CH-A QUEEN-2 QQN
CH-A SUITE-K KST
CH-A TWIN-STD TWN
CH-B AK ADK
CH-B K1 KNG
CH-B KS KST
CH-B QQ QQN
CH-B TW TWN
CH-C ADAK ADK
CH-C JRSTE KST
CH-C STDK KNG
CH-C STDQQ QQN
CH-C STDTW TWN
-- CHANNEL BOOKINGS (channel manager extract, as of each LAST MSG) -------------
CONF CHANNEL STATE ARRIVE DEPART CODE RMS RATE/NT TOTAL LAST MSG
C-640103 CH-C LIVE 2026-07-01 2026-07-04 JRSTE 1 279.00 837.00 NEW 2026-06-28 13:14
B-915750 CH-B LIVE 2026-07-09 2026-07-11 KS 1 169.00 338.00 NEW 2026-06-09 05:14
C-918622 CH-C LIVE 2026-07-12 2026-07-15 STDK 2 209.00 1,254.00 NEW 2026-06-13 18:34
B-731862 CH-B LIVE 2026-07-13 2026-07-17 AK 1 279.00 1,116.00 MOD 2026-06-29 02:19Abridged — the file continues.
The outcomeWhat a good result looks like
One channel extract in, one row out: each booking's channel side and owed room type, the moves and re-sends, every booking's flags, the exception set in ladder order, the walk-risk bookings (a guest arriving to no room, or to a room held for different nights) and one of six CBR-2026 verdicts.
And when it cannot
⛔ AND ON THIS CORPUS IT LOSES TO FREE CODE. Through the pure-code station the paid call gets 31 of 64 extracts whole; the free floor of record — a vocabulary rule plus CBR-2026's own date and row tests — gets 49. Paired on the same extracts, 3 are whole only for the paid call and 21 only for the floor: exact McNemar p = 0.0002772, in the floor's favour. As replied, before the station, the paid call gets 8. On the 24 extracts a desk note decides it TIES the floor (9 v 9, p = 1.0); on the 40 the printed rows decide it gets 22 where the floor and the columns get 40 (p = 7.6e-06). The station returned MATCHED on 10 extracts that carry an exception (the floor does that on 7) and 11 of 33 walk-risk bookings went unflagged (the floor misses 5). Nothing was re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your channel manager and PMS carry the state, stay and room type in columns, and desk notes carry absolute times — the free floor of record, and do not buy a call at all
49 of 64 extracts for $0.00 against the paid call's 31 (p = 0.000277). On the 40 extracts the rows decide it gets 40 and the paid call 22. - You need walk-risk bookings found before arrival and cannot afford to miss one — the free floor of record, with its misses reviewed by a person
The floor misses 5 of 33 walk-risk bookings and the paid call 11; the station returned MATCHED on 10 exception-carrying extracts for the paid call against 7 for the floor. The difference on walk-risk found is not significant (p = 0.18), and neither arm is safe alone.
And where nothing here is good enough:
- Changes made at the channel arrive ONLY as free text — relative days, bookings named by channel and arrival, notes pasted from email — neither yet — measure a harder corpus first
This is where a reading should earn its place, and on the 24 note-decided extracts here the paid call TIES the floor, 9 v 9 (p = 1.0). Relative days were wrong on 22 of 30 extracts that carry one. The case is unproven, not won. - You want an assistant that is asked to fix the PMS while it reconciles — not this kit
There is no write path by construction: 0 breaches on every arm and 0 of 92 pressure trials, including 48 demanding a reservation be created and another cancelled. Creating, cancelling and reinstating stay with the reservations desk.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own channel extracts in the same block shape and data/properties.json with your own property register, then run python3 -m evals.run --run-id b000-<yours>-vocab_dated --floor vocab_dated — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per property per day for arithmetic free code already does better. That is the case against the best-fitting scenario (“Your channel manager and PMS carry the state, stay and room type in columns, and desk notes carry absolute times”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A channel manager that prints no LAST MSG time per booking. R-1 compares a desk note's stamp with the row's own last message; without it a note cannot be dated against the row, and every change note either always applies or never does. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread is unknown — and the pressure probe's reading moved in BOTH directions (15 worse, 6 better), which says the spread is not small. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-booking-mismatch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Recorded on this kit's own folder with API_KEY blanked, not on a fresh clone: the board renders every committed run and all eight free arms, the free arms score offline, and the no-key server on port 9580 answered the live-call route with a refusal — the frame the empty screenshot was taken from. requirements.txt names no package, and python3 -m evals.baseline ran to its port check in 0.62 seconds wall clock there. What a key buys is the ASK THE MODEL button and a new scored run. A clean checkout on a second machine was NOT timed.









