Home › Use Cases › Credential anomaly flagging support
Use caseUC0379
🧪 Use-case kit · runnable

Credential anomaly flagging support

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An admissions or credential-evaluation office receives transcripts from institutions all over the world. Deciding whether one looks ordinary means knowing what ORDINARY looks like for that particular issuer — its grading scale, its credit conventions, its calendar, its course-code format, how it seals and releases a record — and a verifier who does not know an issuer reads its normal conventions as anomalies. That is not a theoretical harm: the cost of getting it wrong lands on an applicant who studied somewhere unfamiliar. the element-by-element comparison a credential evaluator does by eye against an issuer reference file, on the documents they happen to be unsure about rather than on all of them

Audience

The credential evaluator who decides whether to open a verification with the issuing institution, and the head of admissions deciding whether to buy a model call per document. On this corpus the answer is a qualified no on the headline and a qualified yes on one specific thing: the call does not beat free code overall, and it halves the number of legitimate records sent for verification. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual academic transcripts

The corpus is 60 academic transcripts, 0.26 MB (txt 60). EVERY INSTITUTION, COUNTRY, STUDENT, RECORD NUMBER, PROGRAMME AND SEAL IS INVENTED, and on this kit that is not a formality. The output of this system is a list of the parts of somebody's academic record that do not look right; publishing a worked example of that over real documents would be the defect rather than the product. There is no personal data either — a student here is a first name, a family name and a record number, with no date of birth, no nationality and no contact detail. The twelve issuers are given genuinely different-shaped conventions of the kinds that exist in the world — marks out of 20, a scale where 1.0 is best, percentage marks, letter grades on a 10-point scale, modules of 10 to 60 credits, two-, three- and eight-term calendars, five course-code formats — because the whole measurement is what a checker written in one country does to the rest of the world.

The corpus

  • The 60 academic transcriptsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your academic transcripts. That is the whole change — there is no database to migrate.

One academic transcript, as the model receives itTR-0001.txt · 1 of 60
==============================================================================
WEATHERSTONE COLLEGE
Marrow Bay  |  A private liberal-arts college
OFFICIAL TRANSCRIPT OF ACADEMIC RECORD
==============================================================================

--- RECORD ---
Record no.   WEA-2024-16550               Student no.  W245586
Student      Devon Ashgrove
Programme    Bachelor of Arts in Political Science
Issued       05/28/2024                   Issued to    The named student, at their written request.

--- TERMS ---

Fall 2018-2019
  #  Code           Title                              Credits  Grade
  -  -------------  ---------------------------------  -------  -----
  1  ENG 432        Marketing Fundamentals                   1  D+   
  2  BIO 195        Numerical Analysis                       1  B-   
  3  ECON 385       Structural Analysis                      2  C    
  4  BIO 386        Research Design                          4  A-   
  5  ECON 449       Quantitative Methods I                   3  B    

Spring 2018-2019
  #  Code           Title                              Credits  Grade
  -  -------------  ---------------------------------  -------  -----
  6  BIO 344        Digital Systems                          3  C    
  7  POL 378        Strategic Management                     1  B-   
  8  BIO 357        Health Informatics                       2  B-   
  9  MTH 434        Financial Accounting                     4  C-   
 10  ENG 335        Macroeconomic Policy                     3  A-   

Fall 2019-2020
  #  Code           Title                              Credits  Grade
  -  -------------  ---------------------------------  -------  -----
 11  BIO 469        Fluid Mechanics                          3  C-   

Abridged — the file continues.

The outcomeWhat a good result looks like

Every element of every record carries a status, the submitted value quoted as printed, the pattern on file for that issuer, and one of two actions — so a verification request can be written from the report without re-reading the document, and so a reader can check the flag rather than take it.

And when it cannot

It over-flags. 52 elements the key says are perfectly ordinary were sent for human verification — 21.1% of the 246 clean elements — and every one of those is a legitimate record an office would have queried for nothing. It also under-flags: 22 of the 114 genuine anomalies were cleared, and the free pattern table caught every one of them.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A false flag is expensive — you do not want to open a verification against a student who did nothing — the paid call
    52 false positives against the free table floor's 90, 21.1% against 36.6% of the 246 ordinary elements, paired 65 to 27, p = 0.000093. It reads the note that legitimises an unfamiliar convention; the table cannot.
  • A missed anomaly is expensive — you would rather query ten legitimate records than let one through — the free table floor, $0.00
    it flags every genuine anomaly on this corpus by construction, and the paid call's flags are a strict subset of its own: 0 anomalies caught that the floor missed, 22 missed that it caught, p = < 0.000001.
  • You want the whole job and will run one arm — both, because one of them is free
    the floor costs nothing and no network, so running it beside the call is a strictly better position than either alone: the floor's flags bound the recall and the call's flags rank them. Agreement between the two is the cheapest triage signal in the kit.

And where nothing here is good enough:

  • You are tempted to fix the floor's false positives with a keyword scan — neither — measure it first
    that is exactly what table_legend is, and it takes recall from 100% to 12.3%. It clears 0 of the 74 anomalies that carry a near-miss note.

At a glanceHow the whole thing runs

79%flag correct pct
2,015 msp50, end to end
$0.83per 1,000 academic transcripts · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/institutions.json with your own issuer patterns — the id, the grading regex and display, the credit range, the term prefixes, the code format, the seal and date conventions — and drop your transcripts into data/corpus/ with a matching entry in data/transcripts.json. THE MEASURED ACCURACY DOES NOT TRAVEL WITH YOUR CORPUS, and on this kit the gap is likely to be large. Corpus lens →
When is this the wrong choice?Avoid: Do not also expect it to find everything — it misses 22 of 114 genuine anomalies. That is the case against the best-fitting scenario (“A false flag is expensive — you do not want to open a verification against a student who did nothing”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a scanned or photographed credential. Every document here is plain text; reading a sealed PDF or a photograph is a different problem and a different kit, and nothing measured on this page carries over to it. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?whether the model does better against a REAL issuer pattern table. The twelve rows here are authored and internally consistent; a real one is bigger, partly out of date, and disagrees with itself. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-credential-anomaly. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the issuer table, the answer key, the recorded run and all three free floors — and scores every free floor offline. evals/check_labels.py, evals/baseline.py and evals.run --rescore all run with no key, no network and no spend; re-scoring the recorded run reproduces 286 of 360 to the element, because every grader is code. The two commands that need a key are evals.run without a floor and evals.injection, and both say so before they spend.

A living map of modern AI — kept current every morning