Home › Use Cases › Match a scholarship fund's donor terms against its candidate roster
Use caseUC0279
🧪 Use-case kit · runnable

Match a scholarship fund's donor terms against its candidate roster

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A scholarship cycle opens and a financial aid office has a stack of endowed funds, each with its own donor terms written into a gift agreement, and a roster of candidates pulled for each one. Today somebody reads every roster row against every term by hand — a declared major against a named programme, a cumulative GPA against a stated floor, completed credits against a count, a home county against a list — and writes down who is left. The reading is not hard; there is just a great deal of it, and the parts that ARE hard look exactly like the parts that are not. Reading every candidate on a fund's roster against every donor term by hand, and re-reading them when somebody asks why a name is not on the list.

Audience

The financial aid officer preparing a fund's candidate list for the scholarship committee, and the director deciding whether this is a job worth automating at all. The honest answer to the second question is in the floor column: the labelled half is not. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual fund pack

The corpus is 62 fund pack, 0.16 MB (txt 62). A scholarship matching file is the most identifying material a financial aid office holds, and a real one cannot be shipped. Generating it is not a second-best here — it is what lets the answer key be DERIVED rather than typed: every candidate is a set of structured attributes, the pack text is rendered from them, and the key is computed from the same attributes by the same criteria engine every arm goes through. A change to the policy changes the key the moment the generator re-runs. It also lets the hard cases be placed deliberately: 22 of the 62 packs render 74 candidate rows as a registrar's prose note carrying a named trap, each one testing a rule the policy states.

The corpus

  • The 62 fund packgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your fund pack. That is the whole change — there is no database to migrate.

One fund pack, as the model receives itSFM-0001.txt · 1 of 62
SCHOLARSHIP FUND MATCHING PACK
Pack: SFM-0001    Cycle: 2026-27 academic year    Prepared: 2026-03-07
Fund: Ainsley Scholarship for Accounting (fund no. F-1432)
Fund type: endowed, donor-advised terms    Awards this cycle: 1    Award value: $2,500

DONOR TERMS (from the gift agreement, as recorded by Advancement)
  C1  Declared major in Accounting or Agronomy.
  C2  Cumulative GPA of 3.40 or higher.
  C3  Demonstrated financial need: need band high on the current aid file.
  C4  Enrolled full-time.
  C5  Not a prior recipient of an award from this fund.

CANDIDATE ROSTER (Financial Aid pulled 3 records flagged for this fund)
  CAND-01  Student 160199
    Declared major: Accounting
    Cumulative GPA: 3.76
    Class standing: junior
    Credits completed: 95 (in progress: 0)
    Need band: high
    Enrolment: full-time
    Home county: Milwaukee County, WI
    First-generation: yes
    Essay on file: yes
    Prior recipient of this fund: no
  CAND-02  Student 184976
    Declared major: Agronomy
    Cumulative GPA: 3.75
    Class standing: sophomore
    Credits completed: 110 (in progress: 6)
    Need band: high
    Enrolment: full-time
    Home county: La Crosse County, WI
    First-generation: no
    Essay on file: yes
    Prior recipient of this fund: no
  CAND-03  Student 118808
    Declared major: Agronomy
    Cumulative GPA: 3.52
    Class standing: sophomore
    Credits completed: 107 (in progress: 6)
    Need band: low
    Enrolment: full-time
    Home county: Door County, WI
    First-generation: no
    Essay on file: yes
    Prior recipient of this fund: yes (2025-26)

COMMITTEE RECORD
  Scholarship Committee meeting of 2026-02-26.
  Quorum: met (9 of 9 voting members present).

Abridged — the file continues.

The outcomeWhat a good result looks like

Every fund gets a matrix a person can audit cell by cell, a candidate list in roster order, and the roster line behind every term a candidate failed — so a committee argues about the terms rather than about whether somebody read the roster correctly.

And when it cannot

When a roster does not say, the pack must answer NOT-EVIDENCED and carry the candidate as EVIDENCE-INCOMPLETE rather than dropping them. A missing value is not a failing value, and the run's worst behaviour is exactly here: it forced 5 candidate rows the wrong way and on SFM-0058 turned a PENDING-EVIDENCE fund into NO-QUALIFYING-CANDIDATE, taking two students whose file needs one more document off the fund entirely.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your rosters are fixed-layout labelled records — the free rules floor (evals/baseline.py)
    96.1 pct of cells and 70.0 pct of clean packs entirely right, for $0.00 and no network. Read the pack, check the numbers, spend nothing.
  • Your rosters carry prose notes from a registrar — the measured run
    59.1 pct against the floor's 31.8 pct on the 22 hard packs. This is the only place the money buys anything and it is where it buys the most.
  • You need to know which funds have NOBODY qualifying — the measured run, and check the incomplete list by hand
    13 of 13 NO-QUALIFYING-CANDIDATE packs right, and 0 erased into a match even under the adversarial arm. The rules floor answered MATCHED on funds with no qualifying candidate.

And where nothing here is good enough:

  • You want a ranked shortlist for the committee — nothing in this kit
    There is no field in the answer contract that could carry a rank, a score or a name, and the list is scored on ORDER as well as membership. The cap is aid-award, held by the committee.

At a glanceHow the whole thing runs

77%pack all six pct
43,289 msp50, end to end
$26.03per 1,000 fund pack · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace tools/build_corpus.py with something that emits the same three artifacts — one text pack per fund in data/corpus/, a fund register in data/funds.json carrying each fund's criteria and candidates as structure, and data/gold.jsonl. The 77.4 pct does not travel. Corpus lens →
When is this the wrong choice?Avoid: Paying per fund for a job a regular expression already does. The 56.5 pct headline floor is mostly this. That is the case against the best-fitting scenario (“Your rosters are fixed-layout labelled records”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A fund with more candidates than fit in one reply. The whole roster goes in the prompt and the whole matrix comes back cell by cell; the largest reply here was 28,411 output tokens against a 32,000 ceiling on a fund with 8 candidates and 6 terms. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled roster line is the line a scholarship officer would have quoted. The key names one line per failing term and a reviewer might accept a neighbouring line carrying the same fact; this key scores that zero. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-scholarship-match. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — From a clean checkout with NO key configured we reproduced: the corpus rebuild byte-for-byte under two PYTHONHASHSEEDs; the label gate at 0 violations; both free floors scored end to end; the stub run end to end over all 62 packs, proving the whole pipeline including parse, recheck and scoring without a provider; and the local board rendering every pack, both floors and the committed run's replay. What needs a key: the scored run and the adversarial arm, and nothing else. pip install is a no-op — the kit is standard library end to end.

A living map of modern AI — kept current every morning