Home › Use Cases › Special-education transport reconciliation
Use caseUC0367
🧪 Use-case kit · runnable

Special-education transport reconciliation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A student's plan states a transportation service — a service level, a maximum ride time, a day pattern, a monitor or a lift — and a contractor prints a daily run log against it. Whether the entitled service was actually delivered on a given day is not what the STATUS column says: a row can read DELIVERED and be a run made without the one-to-one monitor the plan requires, in a substitute vehicle with no lift, carrying one leg of a two-leg entitlement, or logged against the wrong student; and a row can read NOT-RUN and be a run that did happen, because the cancellation was rescinded or entered elsewhere. The only record of any of it is a sentence printed under the row. On 36 of the 58 files in this corpus the transportation office's own reconciliation panel is wrong, on 24 the run log's own count is wrong, and 19 entitled service days across the corpus were not delivered at all. Opening one reconciliation file, reading the note printed under every run-log row, deciding row by row whether the entitled service was actually delivered, counting the entitled days against the plan's day pattern and the school's excused absences, checking the amendment record against the invoice date, applying the per-diem, netting the other transport against the authorised ceiling and picking one verdict.

Audience

A school district transportation office reconciling a month of routed service against its contractor's invoices, and the special-education administrator who has to know how many entitled service days a child did not get. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation file

The corpus is 58 reconciliation file, 0.26 MB (txt 58). It is generated because a real special-education transport file is a disabled child's education record, their medical needs and their home address in one document, and the exact shapes measured here — a run made without the monitor the plan requires, a missed day nobody reported — are the rows a district would least want published, about the children least able to object. So the corpus contains no student record of any kind, and that is MEASURED rather than asserted: evals/check_labels.py sweeps the shipped bytes of all 58 files against 10 patterns on every run — national identifier, street address, email, telephone, date of birth, disability or diagnosis, family detail, medical detail, behaviour record and a person's name — and reports 0 hits. It ran again while this spec was captured: 0 hits, expected 0. The same module re-derives the whole answer key with its own regexes, importing nothing from src/, at 0 disagreements over 638 checks.

The corpus

  • The 58 reconciliation filegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came from58 files, 268,297 bytes, p50 4,792 and p95 5,037. 968 run-log rows in total, of which 57 are district-closure rows and 137 carry a note; 28 files have something to cite and 32 rows are cited in all; 24 files where the log's own count is wrong and 36 where the office's own panel is wrong; 16 files carrying a real service gap and 19 entitled service days not delivered across the corpus; 4 executed amendments, 12 proposed ones, 12 keep-note decoys, 4 rides over the plan's maximum and 12 files carrying a director's or co-ordinator's instruction note.

Swap this folder for your own material and the kit is pointed at your reconciliation file. That is the whole change — there is no database to migrate.

One reconciliation file, as the model receives itSTR-0001.txt · 1 of 58
==============================================================================
SPECIAL-EDUCATION TRANSPORT RECONCILIATION FILE            STR-0001
District: DIS-2100 - Marchbank Unified School District (invented)
Student: STU-40013   Route: RTE-110   Period: 2026-01-12 to 2026-02-06   Procedure: TSR-2026
==============================================================================

ROUTE AND PERIOD AS THE TRANSPORTATION OFFICE HOLDS IT
  route id                           RTE-110
  contract id                       CON-4420
  contract authorised                    yes
  period authorised                      yes
  service level                 lift-monitor   wheelchair lift and securement, with a one-to-one monitor riding
  day pattern                            M-F
  scheduled service days                  19   in the period, closures excluded
  excused absence days                     2   from the school's attendance record
  entitled service days                   17   scheduled less excused (T-4)
  per diem                            142.00   for the entitled service level (T-5)
  invoice date                    2026-02-17
  gap tolerance                            0   entitled day(s) not delivered
  tolerance amount                      5.00   or less
  tolerance pct                         0.50   pct of supported or less
  period                        school-month

TRANSPORTATION SERVICE AS THE PLAN STATES IT
  entitlement status                IN-FORCE
  effective                       2025-11-16
  service level                 lift-monitor   wheelchair lift and securement, with a one-to-one monitor riding
  monitor                           required
  lift and securement               required

Abridged — the file continues.

The outcomeWhat a good result looks like

One service period in, one row out: which run-log rows this procedure treats differently from the log, each quoted verbatim; whether an executed amendment replaced the printed entitlement status; and then, in pure code, the entitled service days, the delivered days, the missed days, the per-diem total, the ceiling headroom, the supported amount, the unsupported amount and one verdict from a closed set of six.

And when it cannot

And what it does when it cannot. All 58 replies parsed, 0 failed and 0 stopped at the token ceiling, so there is no unanswered row to report. What there IS is over-citation: the arm returned 147 cited rows against the key's 32, and on 13 of 58 periods it reported a service gap the key does not carry. On 3 periods it went the other way and reported a real service gap as RECONCILED, which is the failure that costs a child rather than the district.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your run log's removals all say the same words, and your amendments are decidable from a status column — the free rules floor, and do not buy a call at all
    The floor takes the headline 38 of 58 to 28. It is perfect on the clean files (15 of 15 against 7), perfect on phantom billing (4 of 4 against 0), 12 of 15 against 5 on the rule-only family and 47 of 58 against 26 on the delivered-day count itself. A keyword list over the notes and a comparison on the MIN column reach everything a column can reach, for $0.00.
  • Your findings live in a verb — an amendment that was heard rather than executed, a note about a cancellation of something else — the paid arm
    58 of 58 on the executed-versus-proposed reading against the free regex's 47, which is 11 periods the regex gets wrong in both directions: it applies a merely-proposed amendment on 9 and misses an executed one on 2. TSR-2026's T-6 names missing an executed amendment as the error the rule most exists to prevent, and the arm made it zero times. It is also 5 of 12 against 2 on the proposed-amendment decoys and 57 of 58 against 50 on quoting the row.
  • What you actually need is the count of entitled service days a child did not get, in front of a human — the paid arm, with a second reader on every flagged period
    Of the 19 missed service days in this corpus the transportation office's own reconciliation surfaces 0, the free floor 14 and the paid arm 16. That is the one place the call is clearly worth its money — and it comes with 13 invented gaps on periods where no child missed a run, which is why the second reader is part of the recommendation rather than a footnote.
  • Your finding is a column — a ride over the plan's maximum, an invoice dated before the period, a contract that is not authorised — pure code, and put it in the station rather than in the prompt
    The ride-time cap is decided by the MIN column against the plan's maximum and free code gets the family 3 of 4 against the call's 2; the invoice-date test is 4 of 4 against 3. Everything TSR-2026 decides in code — the entitled day count, the per-diem, the ceiling headroom, both tolerance tests and the verdict ladder — is free on every arm and is published as such.

And where nothing here is good enough:

  • You need a number you can take to a state complaint, a due-process hearing or a parent — neither arm, and not this kit
    TSR-2026 is INVENTED. Its service levels, per-diem schedule, tolerances and rule order are this kit's own and are not anybody's actual obligations to any child, family, district or contractor. Nothing produced here is a determination of entitlement, a compliance finding or a measure of what compensatory service is owed, and the kit refuses in code to produce one.

At a glanceHow the whole thing runs

66%all five correct pct
1,900 msp50, end to end
$0.00per 1,000 reconciliation file · openai/gpt-5-6-luna

Run once, for real, on 2026-09-10. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own reconciliation files in the same printed shape and data/services.json with your own route rows, then re-derive the key with tools/build_corpus.py. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per service period for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your run log's removals all say the same words, and your amendments are decidable from a status column”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A run log whose notes are in a different register — abbreviations, contractor codes or a status vocabulary rather than ordinary English prose. Every reading in this kit rests on the note under the row being a sentence, and both free floors are English keyword lists, so the measurement of what they cost is English too. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One paid run was fired on this corpus, so the run-to-run spread is unknown and unclaimed. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-10 — r002-sped-transport. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run, and both data checks run on a machine with nothing installed: tools/build_corpus.py --check rebuilds the corpus in memory and diffs it against disk, and evals.check_labels re-derives the key and sweeps every shipped byte for personal data. Both were run during this capture and reported 0 problems.

A living map of modern AI — kept current every morning