Home › Use Cases › Check a term's published schedule against the catalog of record
Use caseUC0410
🧪 Use-case kit · runnable

Check a term's published schedule against the catalog of record

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A registrar publishes a term's schedule from what departments request, and the CATALOG OF RECORD is what the institution has actually approved. The two disagree constantly and quietly: a section published with credit hours the catalog does not fix, a restriction copied forward from last year's build, a repeat limit taken off an advising sheet rather than the entry, a cross-listing carried over after the arrangement lapsed. Nothing bounces — registration opens, students enroll, and the disagreement surfaces at graduation audit or at re-accreditation, when it is a transcript problem rather than a schedule problem. The check itself is a clerk reading six published sections against six catalog entries, three prose paragraphs each, in the fortnight before registration opens, for every department. Reading every published section against its catalog entry by eye — the course against the catalog of record, the section number against the rows above it, the credit hours, the grading basis and the mode against the entry's structured half, the scheduled minutes against the contact-minute standard, and then the three prose paragraphs that decide the rest. It does not replace the sign-off; every HOLD goes back to a department and a person owns that conversation.

Audience

A registrar's schedule-build reviewer, reading a department's published sections before the schedule goes live. And the department chair who answers for the rows that come back. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual schedule-consistency packets

The corpus is 62 schedule-consistency packets, 0.26 MB (json 4 · jsonl 1 · md 3 · txt 62). Because the shape of a schedule-consistency failure is not the arithmetic, it is the PROSE. 260 of the 332 sections are supported and 24 of the 72 defects sit in columns any lookup reads. The 48 that matter are sections whose columns are internally perfect and whose catalog entry says something else in a sentence — and a corpus without those measures a parser.

The corpus

  • The 62 schedule-consistency packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 62 packets, data/catalog.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your schedule-consistency packets. That is the whole change — there is no database to migrate.

One schedule-consistency packet, as the model receives itCAT-0001.txt · 1 of 62
SCHEDULE CONSISTENCY PACKET

PACKET HEADER
  Packet             CAT-0001
  Institution        Kestrel Ridge University
  Term               Fall 2026 (term code 202710)
  Catalog of record  2026-2027 Undergraduate Catalog
  Department         Philosophy (PHIL)
  Prepared           2026-04-01
  Prepared by        R. Okonjo, Office of the Registrar

CONTACT-MINUTE STANDARD
  In-person                750 minutes per credit hour per term
  Hybrid                   560 minutes per credit hour per term
  Online (synchronous)     750 minutes per credit hour per term
  Online (asynchronous)      0 minutes per credit hour per term

CATALOG ENTRIES
  PHIL 4663  Introduction to Philosophy
    Credit hours    3
    Grading basis   Pass/fail
    Approved modes  Online (synchronous), Online (asynchronous)
    Registration restrictions
      Prerequisite: PHIL 1661.
      Consent of the instructor overrides the prerequisite at registration.
    Repeatability
      May be repeated for credit up to three times in total.
    Cross-listing
      Listed also as BIOL 3791 and as PSYC 2497.

  PHIL 3429  Quantitative Methods in Philosophy
    Credit hours    4
    Grading basis   Satisfactory/unsatisfactory
    Approved modes  Hybrid, Online (synchronous)
    Registration restrictions
      PHIL 2485L must be carried alongside this course in the same term.
    Repeatability
      A maximum of sixteen credit hours of special-topics work may be applied to the major.
      This course may be repeated for credit to a maximum of twelve credit hours.
    Cross-listing
      Cross-listed with COMM 4386 and with ECON 3722 in alternate years.
      In this catalog year only the COMM 4386 listing is active.

  PHIL 2552  Principles of Philosophy
    Credit hours    3

Abridged — the file continues.

The outcomeWhat a good result looks like

Every published section carries a finding, the rule of CATSCH-2026 it rests on, and one row copied verbatim out of the packet as evidence — the section's own row, the catalog sentence it contradicts, or the contact-minute standard it falls short of. The packet carries PUBLISH or HOLD and the list of rows held. 306 of 332 sections and 11 of 62 packets on the published run.

And when it cannot

⚠︎ THE PACKET-LEVEL NUMBER IS 11 OF 62 AND IT IS WORSE THAN THE FREE COLUMN FLOOR'S 38. That is the single most important sentence on this page. bill_all_correct requires every one of the six graded fields on every row of the packet to be right AND the recommendation and the held list to follow, so one false hold anywhere spoils the packet. The paid call raises 26 false holds across the corpus — supported sections it calls not supported — and they are spread across enough packets to lose the packet-level comparison outright, McNemar exact p = 0.000003. It wins the SECTION-level comparison (306 against 284, p = 0.014080) because it finds all 48 defects a lookup cannot reach. This is a finder, not a gate.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your departments code restrictions, repeat limits and cross-listings correctly, and your schedule-build disputes are about credit hours, grading bases, modes and contact minutes. — the free column floor alone — python3 -m evals.run --floor rules
    Every one of those is a lookup, a comparison or one multiplication. The floor is 284 of 332 sections and 38 of 62 packets, deterministic, and $0.00.
  • Your catalog states restrictions, repeat limits and cross-listings in prose — which every catalog does — and your build copies them forward by hand. — the paid call, rechecked
    Those are the 48 sections the column floor is wrong on by construction, and the call gets 48 of 48 of them.
  • You were going to write a regex over the catalog prose instead. — read the phrase floor's number first
    It is on this page: 139 of 332 sections, worse than reading no prose at all, because a misread paragraph turns a supported section into a false finding.

And where nothing here is good enough:

  • You want one number to gate the schedule build on. — neither, on this evidence
    The free column floor wins the packet-level comparison 38 to 11 (p = 0.000003) and the paid call wins the section-level one. Neither is a gate; the call is a finder and a person signs the build off.

At a glanceHow the whole thing runs

92%section finding correct rechecked pct
2,822 msp50, end to end
$1.88per 1,000 schedule-consistency packets · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported packets and data/catalog.json at your own catalog records. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need, once per department per term. That is the case against the best-fitting scenario (“Your departments code restrictions, repeat limits and cross-listings correctly, and your schedule-build disputes are about credit hours, grading bases, modes and contact minutes.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose panels are not the seven this parser knows. src/rules.py splits on PACKET HEADER, CONTACT-MINUTE STANDARD, CATALOG ENTRIES, PUBLISHED SECTIONS, SCHEDULE BUILD NOTES, SIGN-OFF and END OF PACKET; a missing heading yields an empty section rather than an exception, so a differently-shaped packet parses to zero rows and is scored as zero rows. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from a real difference. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-catalog-schedule. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all four floors score all 62 packets in under a second, and python3 -m evals.check_labels re-derives the whole key from a hand-retyped rulebook. Both are $0.00.

A living map of modern AI — kept current every morning