The business caseThe problem this solves
A school's special-education scheduling desk drafts a team meeting notice for every meeting request: the date and start time, the room, the purpose in the district's fixed wording, every invited role in order, the family's language and a reply-by date. Almost all of it is printed on the request packet. What is NOT printed on any column is whether a printed value may still be used: a room taken out of use for floor works, a case manager going on leave, a purpose code under revisit, a family contact's language being corrected — each with an effective date, each withdrawable by a later entry, each sometimes naming the held twin rather than the booked record. All of that arrives as sentences in the scheduling log, and today a coordinator reads it by eye before the notice goes out. Reading one meeting notice request packet's scheduling log against the district's own notice template to decide whether each of the four elements can still be used, then filling in the notice by hand — not the decision to send it to a family, and not the coordinator who owns it.
Audience
A special-education coordinator deciding whether a drafted notice can go to review, and the person who owns the district's scheduling tooling deciding whether a model call belongs in that path. The answer this report gives them is NO, with the measurement attached: on this corpus the paid call is lower than free rules-and-regex code that reads the log, and beats only code that does not read the log at all. What it does buy is measured too: it declines significantly more of the notices the template forbids, and it declines significantly more notices nobody forbade. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual meeting notice request packets
The corpus is 64 meeting notice request packets, 0.19 MB (json 5 · jsonl 1 · md 2 · txt 64). Because the failure that matters on a meeting notice is not filling in fields, it is using a value the log has quietly put out of use — and the log is prose. Every packet carries six entries: deciding ones, decoys that name the held twin, entries dated after the comparand, entries later withdrawn by name, a withdrawal, and a benign line. A reader that spots the word 'closure' and stops is measurably worse than one that never looks: the domain word list takes 22 and the generator-phrase regex without structure 19, against the log-blind columns' 25. The draftable/declined mix (38/26) and the 16 packets with no window supplied are fixed by the seed and reported, not engineered away.
The corpus
- The 64 meeting notice request packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 64 packets, data/request_records.json and the whole key are generated in-process by the file that renders them, so there is no third-party data and no third-party licence.
Swap this folder for your own material and the kit is pointed at your meeting notice request packets. That is the whole change — there is no database to migrate.
TALLOWBROOK VALLEY SCHOOL DISTRICT - TEAM MEETING NOTICE REQUEST PACKET
SYNTHETIC SAMPLE DATA. GENERATED, NOT REAL. NO REAL DISTRICT, SCHOOL, STUDENT OR FAMILY.
Packet MNR-0001 Request REQ-48207 Prepared 2026-04-02
CASE
Student reference STU-37427
School Kestrel Park Elementary
Grade 3
Request received 2026-03-29
Purpose code evaluation-results
Previous request on this case REQ-48763
MEETING ON FILE
Booked slot SLOT-4677
Booked meeting date 2026-04-12
Booked start time 13:15
Booked room KP-321
Chair (case manager) CM-2604
ALTERNATE SLOT HELD, NOT BOOKED
Held slot SLOT-4586
Held slot date 2026-04-18
Held slot start time 16:15
Held room KP-295
Held slot chair CM-2448
FAMILY CONTACTS
Primary contact for notices FAM-51248
Primary contact's language of communication Vietnamese
Second contact on record FAM-51254
Second contact's language of communication Portuguese
INVITED ROLES (in the order the notice lists them)
1 Parent or guardian
2 Case manager
3 General education teacher
4 Special education teacher
5 District representative
6 School psychologist
7 Speech-language pathologist
NOTICE TEMPLATES ON FILEAbridged — the file continues.
The outcomeWhat a good result looks like
Either a drafted notice of seven fields — meeting date, start time, room, purpose, every invited role in printed order, the family's language, reply-by date — routed to COORDINATOR-REVIEW, or exactly ONE named element (meeting-slot, location, purpose or family-language) that is absent or beyond use, routed to ELEMENT-CHASE with no notice fields at all. Every packet also carries the template in force and a timing verdict against the OPERATOR-SUPPLIED lead-time window, or TIMING-UNKNOWN when none is supplied.
And when it cannot
⛔ THIS KIT DOES NOT BEAT FREE CODE, AND THE HEADLINE IS PUBLISHED AS A LOSS. The measure is the FULL ROW — the decision, the named element, the template, the timing, the route and all seven notice fields, on one packet. After the pure-code station the paid call takes 45 of 64. The free floor of record — a frozen domain word list plus the record join, the effective-date test and keyed withdrawals, $0.00 — takes 50. Paired on the same 64 packets that is 7 against 12 discordant, McNemar exact p = 0.359283: not significant, and the paid call is the LOWER arm. Before the station the paid call's own answer is 32 of 64, and that loss IS significant (p = 0.001431). It beats only code that never reads the scheduling log: 45 against 25, p = 0.000821.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A district that can write its template's rules and log wordings out in code — the free rules/regex floor, and no model call
it takes 50 of 64 full rows at $0.00 and no network; the paid call takes 45 and the paired test cannot separate them (p = 0.359283) — with the paid call behind. - A log whose wordings vary beyond any fixed vocabulary — the paid call with the station, as a second reader beside the free floor — not instead of it
on this corpus the two disagree on 19 packets: the paid call is right on 7 the floor misses (a closure worded outside the frozen list), the floor right on 12 the paid call misses. - The harm you are guarding against is a notice that should never have gone out — the paid call with the station, as a SECOND reader whose declines a coordinator checks
on the 26 packets the template says must be declined it declines 21 against the floor's 12 (p = 0.011719): the floor drafts 14 notices it should not, the call 5. - No lead-time window configured — any arm — timing is TIMING-UNKNOWN by rule
every code reader holds TIMING-UNKNOWN on all 16 no-window packets; the paid call's own answer invented a verdict on 1 (MNR-0028), which the station then overwrote.
And where nothing here is good enough:
- Anything a family will read — neither; a coordinator
5 packets got a drafted notice the template forbids, on both columns, and the station cannot see them because it trusts the reading.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered request packets and data/request_records.json at the structured half, rewrite src/packet.py for your layout, and replace data/notice-template.md + data/notice-template.json with your district's own template. The boundary is the ANSWER KEY. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying a call to beat a floor it does not beat on this corpus. That is the case against the best-fitting scenario (“A district that can write its template's rules and log wordings out in code”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet not laid out the way src/packet.py parses it. The parser is positional over one rendering; another district's scheduling export is a new parser, not a new prompt. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | It does not beat free code, and that is measured rather than unverified: 45 of 64 full rows against the rules/regex floor's 50. What could not be verified is whether a larger corpus would move it — 64 paired packets give p = 0.359283, and no bigger corpus was bought. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier — the paid call, after the station. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-meeting-notice-draft. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Recorded on a clean checkout with no key configured and no network: python3 -m evals.baseline scores the free readers over all 64 packets and writes the committed b000 result files; python3 -m evals.check_labels re-derives the key and reports 0 disagreements; python3 -m evals.run --selftest passes with no socket; and the board serves every panel from results/ with the one spending endpoint returning 200 skipped and its reason. A clean checkout cannot buy a call — that needs one credential — and nothing on the free path opens a socket.









