Home › Use Cases › Read a hearing decision's ordered relief and say what the case record evidences
Use caseUC0517
🧪 Use-case kit · runnable

Read a hearing decision's ordered relief and say what the case record evidences

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A hearing decision is issued, it orders a handful of things, and a month later an implementation unit has to answer one question about it: of the things this decision ordered, which does the case record actually evidence? The decision does not mark its orders. A trailing conditional opens exactly like an order and ends on a condition that may never fire; an admonition is written with "shall"; an operative instruction sits in the findings rather than the Order; a restatement says the same thing twice and is one obligation, not two. On this corpus 88 sentences read like orders and are not, and 47 of 64 decisions carry a compound sentence that orders two things in one breath. "Was this decision issued" is a column somebody typed and needs no reader. an implementation-unit spreadsheet whose row per decision is somebody's retyped summary of what the decision ordered — which merges a compound disposition into one line, splits a restatement into two, quietly carries a sentence that only looked like an order, and loses the entry from the previous review that is the only evidence half the obligations have.

Audience

an implementation-unit reviewer preparing an open hearing decision for its monthly read, and the hearings coordinator who will make the findings this kit refuses to make. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual hearing decisions under implementation review

The corpus is 64 hearing decisions under implementation review, 0.48 MB (txt 64). There is no public corpus of fair-hearing decisions paired with the case record entered against them afterwards, and there will not be: the pairing is the private half. A redacted decision alone is worse than useless here — what this kit measures is the join between an order and an entry, and redaction removes exactly the entry. So the corpus is invented end to end, and data/SOURCES.md carries a section on what the generator costs the measurement: the sentence-opener tell that had to be designed out (orders and non-orders now share at least two openers, 7 shared across 47 of 149 ordered sentences), the compliance window placed leading, trailing and absent so "strip after the last comma" is not a way past the refusal, and a lure noun that collided with an artefact noun and made 5 gold rows unanswerable until it was found by asking why the ceiling arm stopped at 99.08%.

The corpus

  • The 64 hearing decisions under implementation reviewgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your hearing decisions under implementation review. That is the whole change — there is no database to migrate.

One hearing decisions under implementation review, as the model receives itHD-0001.txt · 1 of 64
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated fair-hearing decision and post-decision
case record for an AI use-case kit; it reproduces no agency, no household, no real person,
no real case and no real decision. The implementation standard is an ILLUSTRATIVE,
OPERATOR-TUNABLE default, not any agency's standard and not a statute or a regulation.

Decision Caption
----------------------------------------------------------------
Decision reference                 : HD-0001
Docket                             : FH-26-0107
Case reference                     : C-88-1013
Appellant                          : the household of Vasquez
Program at issue                   : child care assistance
Decision issued                    : 2026-01-24
Decided by                         : a hearing officer of the agency's appeals unit
Certification period at issue      : certification period P-2203
Other certification period on file : certification period P-2250  (not in issue)

Issue Presented
----------------------------------------------------------------
Whether the Department's action on the household's child care assistance allotment for
certification period P-2203 can stand on the record before the hearing officer.

Findings Of Fact
----------------------------------------------------------------
Every sentence in this section, in Conclusions and in the Order carries a sentence id in the
margin. An id is a citation device only: it says nothing about whether the sentence orders
anything.

  S01  The household of Vasquez appealed the Department's action on case C-88-1013. The
       household's child care assistance certification at issue is certification period

Abridged — the file continues.

The outcomeWhat a good result looks like

the exact set of ordered actions with the sentences they are stated in, the sentences that read like orders and are not, and an evidence state per obligation with the one case-record entry it rests on — 35 of 64 decisions enumerated exactly against a free-arm floor of 1 of 64.

And when it cannot

a decision that orders nothing returns an empty obligations list, which is a real answer and is scored as one — 2 of the 64 are like that. An obligation whose record carries nothing relating to it is NONE with a null record line; no entry is invented to fill it, and a row that does not exist is never scored as a correct null.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • The unit already keeps a clean obligation list per decision, typed by a reviewer, and only wants the evidence state filled in — the free floor -- domain_noun, $0.00
    Of the rows it reaches it calls the evidence state right 80.27% of the time against the paid arm's 69.63%. Handed the enumeration it does not have to find it.
  • The list has to be derived from the decision itself, compound sentences, restatements and trailing conditionals included — the call -- and buy it for the enumeration and the non-order call, nothing else
    35 of 64 decisions enumerated exactly against 1 of 64, p < 1e-8, and 75 of 88 non-order kinds right against 0 of 88 on every free arm.
  • The output must never state a date, a lapse or a compliance finding, and somebody has to be able to prove it — the call, with evals/align.py::refusal_violations in the loop
    0 violations on 64 of 64 decisions and 24 of 24 adversarial trials, against 123 violations over 41 of 64 for free code that quotes the page back, p = 9.10e-13.

And where nothing here is good enough:

  • You want a compliance finding, a due date or a lapse flag — neither, and not this kit
    There is no field in the answer contract in which one could be written, and the grader convicts any attempt. That gate belongs to the hearings coordinator.

At a glanceHow the whole thing runs

55%set exact pct
1,421 msp50, end to end
$1.97per 1,000 hearing decisions · Gemini 3 Flash

Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?tools/build_corpus.py writes data/corpus/*.txt, data/gold.jsonl and data/corpus-stats.json. The measured figures on this page do not transfer to your own decisions. Corpus lens →
When is this the wrong choice?Avoid: Do not use it to build the list. It enumerates 1 of 64 decisions exactly, merges 40 rows and emits 58 the key has nothing for. That is the case against the best-fitting scenario (“The unit already keeps a clean obligation list per decision, typed by a reviewer, and only wants the evidence state filled in”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A decision whose order section is a 40-paragraph consent schedule, or a case record with several hundred entries. One decision must fit in one call; there is no chunking step and no retrieval step to fall back on. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether a second tier reproduces either the enumeration result or the three dangerous false positives. One tier, one run, one corpus. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-21 — r001-ordered-relief. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, replays all 64 committed replies, computes the strongest free arm live in pure Python and shows the answer key beside both. python3 -m evals.check_labels runs its 37 assertions and python3 -m evals.baseline regenerates all six free arms, both offline and both at $0.00. What a clean checkout cannot do without a key is fire a new scored run: the read button on the board reports that in a plain sentence rather than an error.

A living map of modern AI — kept current every morning