Catch buried assumptions in a mortgage lender's appraisal reports
Appraisal reports can hide an assumption about a septic system or a foundation mid-paragraph, with no heading. This app reads each report, fills in its key details and sends any report that states an assumption to a review appraiser.
PresenterOpens the private repo. Visible to admins only.
For the appraisal review deskCross-domain · Banking
Why it matters
Today's manual process, and the same job with the app
A collateral review team at a mortgage or consumer lender, checking appraisal reports before a loan decision.
✕Today's manual process
1Read every report end to end, not just the Extraordinary Assumptions heading.
2Copy the details into the loan file: address, appraiser, dates, value and square footage.
3Hunt for assumptions buried mid-paragraph in Scope of Work or Comments, with no heading.
4One missed assumption means a loan decided on an appraisal that should not be relied on.
Every report read in full, manually
✓With the app
1Every report is read in full, every section, not only the headed ones.
2Ten details are filled in, most showing the section of the report they came from.
3Any stated assumption is pulled out word for word, wherever it sits.
4Those reports go to a review appraiser, who still makes the call. Nothing is accepted or rejected automatically.
Flagged reports go straight to review
See it work
One real case, read by the app, step by step
Avery C. Duarte's report on 8810 Willow Creek Rd assumes a workmanlike completion, in a line buried under Scope of Work.
Catch buried assumptions in a mortgage lender's appraisal reportsReference appBuilt to be shaped to your process
5
1The report's details address, appraiser, both dates, the approach and the reconciled value from one report.
2Where each came from the section of the report behind a detail, so it can be checked.
3An assumption is stated yes, even though the report has no heading for one.
4The buried line a workmanlike completion is assumed, found under Scope of Work.
5Sent for review a review appraiser decides; the app never accepts or rejects the value.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch buried assumptions in a mortgage lender's appraisal reports
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Checking an appraisal report for a stated extraordinary assumption means reading the whole report, not just a section labelled Extraordinary Assumptions -- an assumption about the septic system, the foundation or an environmental condition can be embedded mid-paragraph in Scope of Work or Comments prose with no heading marking it as special, and missing one changes whether the appraisal should be relied on for the loan decision. Someone manually reading each appraisal report end to end to check whether it states an extraordinary assumption anywhere — not just under a dedicated heading — before the file goes to a review appraiser.
Audience
Collateral/appraisal reviewers in mortgage and consumer lending, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reports
The corpus is 55 reports, 0.04 MB (txt 55). Plain text, one format, invented rather than fetched — a real appraisal report is not something this repo can publish either (see SOURCES.md). Ten fields are chosen because the same ones matter to a collateral/appraisal review: who, where, the two dates, the approach, the reconciled value, and — the field this kit exists to test — whether an extraordinary assumption is stated anywhere in the report.
The corpus
The 55 reportsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your reports. That is the whole change — there is no database to migrate.
One report, as the model receives itAR-0001.txt · 1 of 55
Subject Property
----------------
7470 Willow Creek Rd, Clearview, CO
Appraiser
---------
Casey D. Rossi
Effective Date
--------------
2026-06-06
Report Date
-----------
2026-06-09
Valuation Approach
------------------
cost
Improvements
------------
Gross living area: 4011 sq ft
Comparable sales used: 5
Scope of Work
-------------
The appraiser inspected the subject property and researched the local market for comparable sales. This report was prepared in conformance with USPAP.
Comments
--------
No further comments.
Extraordinary Assumptions
-------------------------
It is assumed the subject property will be completed in a workmanlike manner per the plans and specifications provided, as the improvements were under construction at inspection.
Reconciliation
--------------
Reconciled value opinion: $318,700
The outcomeWhat a good result looks like
A ten-field extracted record per report, plus one pure-code computed routing flag -- whether the report states an extraordinary assumption anywhere -- that routes it to a certified review appraiser. Never an automated USPAP compliance judgment or value acceptance.
And when it cannot
This run found zero errors on the safety-critical field (extraordinary_assumption_present, 55/55 correct on both tiers) -- there is no observed flag failure to report from the run itself. What the run did find: extraordinary_assumption_text narrows to the core assumption sentence rather than gold's full surrounding paragraph on 10 of 29 stated cases, identically on both tiers -- a text-boundary choice that never changed presence detection or the review flag. What the run did not test: a report whose extraordinary assumption is stated only in an addendum the main body references, or an adversarially-worded report -- see Data.breaks_on and Eval.could_not_verify.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Screening a stack of appraisal reports for which ones need a certified review appraiser to look at an extraordinary assumption before the file closes — either tier for the flag — they tie at 1.00 recall and precision measured here 100% EA-present recall on both tiers, against the free heading-only floor's 62.07% recall (11 missed review flags, all of them embedded assumptions with no dedicated heading).
Deciding the two tiers on cost, speed or latency variance — the fast tier Roughly half the p50 latency (2,366 ms vs 4,823 ms), a p95 nearly 6x lower (5,872 ms vs 34,729 ms, with one deliberating-tier call taking 154 seconds), and 51% cheaper per query (see Cost) for identical extraction and flag scores.
At a glanceHow the whole thing runs
98%extraction accuracy
2,366 msp50, end to end
$1.39per 1,000 reports · Google Gemini 3 Flash
Run once, for real, on 2026-08-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch buried assumptions in a mortgage lender's appraisal reports14 steps · 4 questions · run once, for real · 2026-08-21
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per report with a stated flag for the optional text field. Corpus lens →
When is this the wrong choice?
Avoid: The heading-only floor for anything beyond a labelled Extraordinary Assumptions section — it is exactly the shortcut this corpus's planted ambiguity is built to defeat. That is the case against the best-fitting scenario (“Screening a stack of appraisal reports for which ones need a certified review appraiser to look at an extraordinary assumption before the file closes”). 2 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Scanned or image-only reports — there is no OCR step; this corpus assumes a text-extractable report, per this kit's own facet sheet. 3 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Why the deliberating tier's output tokens and tail latency were so much higher on this corpus than the 8-9% gap this series' sibling kits saw is unexplained. The gap here: 572 tokens/call average versus 336, one call running 154 seconds. 3 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the deliberating tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-21 — r001-appraisal-extract. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Verified from a cold clone with no key configured: 55 reports segment into 504 sections in 0.002 seconds, and the assembled prompt for AR-0003 replays byte for byte — its three part sizes match what run r001 recorded.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
2,366 msp50, end to end
5,872 msp95
2 minclone to first result
What the clock covers. model call only, one per report
Current processWhat it replaces
Someone manually reading each appraisal report end to end to check whether it states an extraordinary assumption anywhere — not just under a dedicated heading — before the file goes to a review appraiser.
Where it is not good enough
It found every stated extraordinary assumption on both tiers (29/29, 1.00 recall and precision) — the safety-critical figure — but its verbatim extraordinary_assumption_text sometimes narrows to the core assumption sentence rather than the full surrounding paragraph gold recorded, on 10 of 29 stated cases on both tiers identically. That never affected presence detection or the review flag, but it is a real field-level accuracy gap on the one field besides the flag that matters most here — see Eval.taxonomy.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free heading-only floor (evals/baseline.py, checks ONLY a section literally titled "Extraordinary Assumptions") scores 95.8 pct extraction accuracy but 62.07 pct flag recall — 11 missed review flags, every one an assumption embedded with no dedicated heading — against 100 pct flag recall and precision on both tiers here, with zero flag errors of any kind. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
SECTION_HINTS
src/select.py
map fields to your own report's headings; unmatched falls back to the whole document
PROVIDERS
src/adapters/__init__.py
any OpenAI-compatible host, or Anthropic's Messages API
the field schema
data/fields.json
a different set of fields entirely, with its own types and allowed values
Components
Component
File
Role
segment
src/segment.py
cut the report into addressable sections, pure code
select
src/select.py
pick which sections carry each field, pure code — three sections for the extraordinary-assumption fields, since it can be stated in any of them
prompt
src/prompt.py
assemble one call for all ten fields
extract
src/extract.py
the AI layer, one provider one key — plus the pure-code computation downstream: the review-routing flag
judge
evals/judge.py
score field accuracy and the flag's recall/precision separately, pure code
Where it breaks at scale
One call per report, no concurrency and nothing shared between calls: 55 reports took a few minutes wall clock on the fast tier — the deliberating tier's p95 ran to 34.7 seconds on one report, a real tail this kit does not smooth over. A stack of thousands needs batching and a rate-limit strategy this kit does not have.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Before anything is asked of the model. Ten named fields with their own types and allowed values.successOpen full size →AR-0003, extracted live. An extraordinary assumption about the septic system, embedded mid-paragraph in Scope of Work with no dedicated heading, correctly reads extraordinary_assumption_present: yes — the source span names § Scope of Work, not a labelled Extraordinary Assumptions section that does not exist on this report — and the computed panel routes it: "YES — route for review appraiser."successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same button with no API_KEY configured. A calm 200 and a plain sentence, not a stack trace: nothing was called, nothing was spent, and the field table stays browsable.failureOpen full size →
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
55reports
0.04 MiBtxt 55
504sections · p50 56 chars
$0.00setup · 0.002s
How it is cutWhat one section is
cut on underlined section headings; a report with none falls back to one whole-document segment so a span still resolves
SetupWhat the setup figure measured
There is no index. Preparation is segmentation only — 55 reports cut into 504 sections by src/segment.py, pure code, no model and no key.
LicenceLicence
MIT — this repository's own licence. Every property, appraiser and dollar figure is invented; a real appraisal report cannot be published either — it carries a real address and a real licensed appraiser's identity tied to a real client's loan file.
Bring your ownBring your own reports
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per report with a stated flag for the optional text field. SECTION_HINTS in src/select.py maps fields to headings and will need editing for a different report layout; when it does not match, selection falls back to the whole document — slower, more expensive, always correct.
What breaks it
Scanned or image-only reports — there is no OCR step; this corpus assumes a text-extractable report, per this kit's own facet sheet.
A report whose sections are not headed — segment() falls back to one whole-document segment, so a span names "document" and locates nothing finer.
An extraordinary assumption stated only in an addendum the main body references but does not repeat — this corpus never places one there.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
934
286
field schema
1,130
265
report sections
719
228
Total
779
This is the cost lesson as arithmetic: of the 779 tokens assembled, 286 are instructions — 37% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Measured token-for-token via evals/prompt_tokens.py's nested-prefix subtraction against the provider's own tokenizer, not estimated from characters — see results/tokens-p001-appraisal-extract.json.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
SYSTEM:
You extract structured fields from a real-estate appraisal report. You return JSON and nothing else.
RULES, in order of importance:
1. If the report does not state a field, return null for it. Do not infer it, do not compute it, and do not use what you know about the world.
2. `extraordinary_assumption_present` means the report states an extraordinary assumption SOMEWHERE in its text -- this may be under a dedicated 'Extraordinary Assumptions' heading, OR embedded in the prose of a different section (Scope of Work, Comments, or elsewhere) with no special heading at all. Read the entire report before answering 'no'; a report with no dedicated Extraordinary Assumptions section may still contain one embedded in another section's text.
3. Copy values verbatim from the report wherever possible.
4. Use the exact allowed value for a field that lists them.
5. Return every field named in the schema, even when the answer is null.
USER:
Extract these fields:
- property_address (string) -- the subject property's address, verbatim
- appraiser_name (string) -- the appraiser's name
- effective_date (string) -- the appraisal's effective (value) date, YYYY-MM-DD
- report_date (string) -- the date the report was written, YYYY-MM-DD
- approach_used (enum) one of: sales_comparison, cost, income -- the valuation approach stated on the report
- reconciled_value (number) -- the final reconciled value opinion, no currency symbol or commas
- gross_living_area_sqft (integer) -- the subject's gross living area in square feet
- comparable_count (integer) -- how many comparable sales the report states it used
- extraordinary_assumption_present (enum) one of: yes, no -- does the report state an extraordinary assumption ANYWHERE in its text -- under a dedicated 'Extraordinary Assumptions' heading, OR embedded in the prose of another section (Scope of Work, Comments, etc.)? Read the whole report; do not check only a labelled section.
- extraordinary_assumption_text (string) -- the extraordinary assumption's text, copied verbatim from wherever it appears, or null if the report states none
Return a JSON object with exactly these keys: property_address, appraiser_name, effective_date, report_date, approach_used, reconciled_value, gross_living_area_sqft, comparable_count, extraordinary_assumption_present, extraordinary_assumption_text
Use null for any field the report does not state.
REPORT
------
Subject Property
----------------
1408 Stonegate Ct, Fairmont, NC
Appraiser
---------
Jordan O. Chen
Effective Date
--------------
2026-03-18
Report Date
-----------
2026-03-19
Valuation Approach
------------------
sales comparison
Improvements
------------
Gross living area: 3172 sq ft
Comparable sales used: 4
Scope of Work
-------------
The appraiser was not provided a current survey; the site boundaries and improvements are as represented by the county assessor's records. It is assumed the subject's septic system is functioning adequately; no inspection was performed and none was available for review. This report proceeds on that basis.
Reconciliation
--------------
Reconciled value opinion: $678,400
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"property_address": "1408 Stonegate Ct, Fairmont, NC",
"appraiser_name": "Jordan O. Chen",
"effective_date": "2026-03-18",
"report_date": "2026-03-19",
"approach_used": "sales_comparison",
"reconciled_value": 678400,
"gross_living_area_sqft": 3172,
"comparable_count": 4,
"extraordinary_assumption_present": "yes",
"extraordinary_assumption_text": "It is assumed the subject's septic system is functioning adequately; no inspection was performed and none was available for review. This report proceeds on that basis."
}
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch buried assumptions in a mortgage lender's appraisal reports — 55 reports. Two tiers of one model family answered, and every answer was then graded Two different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
55reports
55source documents
2model tiers
110graded answers
2grading methods
MeasurementsWhat was measured
COUNTED514 · 514 / 524extraction accuracy — stated cellsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED29 · 29 / 29EA-present recall — reports that state an extraordinary assumptionDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/judge.py) is pure code, exact match with light normalisation against a mechanically-derived gold — there is no judgement to validate, only comparison. What WAS validated: tools/build_corpus.py's gold is read back off the same generated report text, never carried over from a separate target, and evals/check_labels.py asserts stated/value and present/text consistency on every row before any run is allowed to spend. The TRUE review flag is derived by running the same pure-code compute() over gold's own values, never a separately-typed truth.
335.8output tokens · the fast tier · 2,366 ms p50
572.15output tokens · the deliberating tier · 4,823 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 2.0× as long, and lands one row apart on 55. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One report
1,000 reports
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.001395
$1.39
28%
Same work, 1× the bill
The same reports, the same tokens — only the rate card changed. And on that card about 28% of what you pay is the prompt this pipeline sends, not the answer it writes.
which tier is called — the two tiers tie on every accuracy and recall figure measured here, so the lever buys accuracy nothing; it does buy 51% lower cost per query and a far tighter latency tail on the fast tier.
Rates checked 2026-08-18. The provider that actually ran r001 and r002 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page.
The gradersTwo ways to grade
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Per-field exact match, light normalisation Does the model's value for each field match gold, after trimming whitespace/punctuation and treating numbers within half a cent as equal?
$0.00
no
yes
the fast tier 98.1% · the deliberating tier 98.1%
Review-flag confusion matrix Does the run's own pure-code needs_review (computed from what the model extracted for extraordinary_assumption_present) match the same computation run over gold's own true values?
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 100.0%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, between the models and the heading-only floor, not between the two model tiers. Both tiers hit 55/55 on extraordinary_assumption_present (the safety-critical field) and matched each other cell-for-cell on every other field, including the identical 10-of-29 boundary gap on extraordinary_assumption_text (see Eval.taxonomy). The floor is what separates cleanly: 62.07% recall against 100%, on exactly the embedded cases this corpus was built to plant.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Screening a stack of appraisal reports for which ones need a certified review appraiser to look at an extraordinary assumption before the file closes
either tier for the flag — they tie at 1.00 recall and precision measured here
100% EA-present recall on both tiers, against the free heading-only floor's 62.07% recall (11 missed review flags, all of them embedded assumptions with no dedicated heading).
the heading-only floor for anything beyond a labelled Extraordinary Assumptions section — it is exactly the shortcut this corpus's planted ambiguity is built to defeat.
Deciding the two tiers on cost, speed or latency variance
the fast tier
Roughly half the p50 latency (2,366 ms vs 4,823 ms), a p95 nearly 6x lower (5,872 ms vs 34,729 ms, with one deliberating-tier call taking 154 seconds), and 51% cheaper per query (see Cost) for identical extraction and flag scores.
paying for the higher tier here — it bought nothing measurable and cost meaningfully more, with a real latency tail the fast tier did not show.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
text-boundary-narrowed
Extraordinary-assumption text extracted as the core sentence only, not gold's full surrounding paragraph
10
AR-0003: gold's extraordinary_assumption_text is the full three-sentence embedding ("The appraiser was not provided a current survey...It is assumed the subject's septic system is functioning adequately...This report proceeds on that basis."). Both tiers…
What we could NOT verify
Why the deliberating tier's output tokens and tail latency were so much higher on this corpus than the 8-9% gap this series' sibling kits saw is unexplained. The gap here: 572 tokens/call average versus 336, one call running 154 seconds. No thinking parameter was sent on any call (see LLM.settings), so this is provider-default behaviour, not a setting this kit's own harness controls, and it was not diagnosed further.
Whether the model would still hit 100% EA-present recall on a larger, adversarially-constructed set of embedded assumptions — this run's 11 embedded cases all resolved correctly on both tiers, but 11 cases is not enough to rule out a harder placement this corpus did not think to plant.
How either tier performs on a real, messy appraisal report layout (multi-page PDFs with addenda, form-based URAR layouts, scanned images) — this corpus is plain text with a single, consistent underlined-heading layout by construction.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
774.62
335.8
2,366 ms
$0.001395
the deliberating tier
774.62
572.15
4,823 ms
$0.002104
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The heading-only floor (evals/baseline.py) and the scorer (evals/judge.py) are both pure code and cost $0.00 to run against either result set. The figure above is both runs' own token counts (the fast-tier and deliberating-tier runs, 55 reports each) priced at Google Gemini 3 Flash's published rate -- the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The fixed system prompt and field schema (551 of 779 tokens on the example call, 71%) outweigh the report sections sent (228 tokens) on this example — the floor every call pays before a single report fact is read.
Output length on the deliberating tier: 572 tokens/call average against the fast tier's 336 for an identical ten-field JSON record, plus a heavy tail (one call at 154 seconds) — see Eval.could_not_verify for what is and is not known about the cause.
Your volumeWhat it costs at your volume
Linear in reports: each call is independent and self-contained, with no shared context or index to amortise. This run's 55 reports cost about $0.077 projected onto Gemini 3 Flash's rate, so ten times the set is about $0.77 on the same rate and the same prompt — arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
42,604input tokens · this run
18,469output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 55 appraisal reports extracted, each scored by pure code against a mechanically-derived gold set. This run (r001-appraisal-extract, the fast tier) answered all 55 of 55 reports with no truncation -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.031
$0.031
$0.56
2026-09-12
gemini-3-flash
Google
$0.077
$0.077
$1.39
2026-09-18
gemini-3-8-flash
Google
$0.101
$0.101
$1.84
2026-09-18
llama-5
Meta
$0.132
$0.132
$2.40
2026-09-18
claude-haiku-4-5
Anthropic
$0.135
$0.135
$2.45
2026-09-12
grok-4-5
xAI
$0.196
$0.196
$3.56
2026-09-18
grok-4-6
xAI
$0.196
$0.196
$3.56
2026-09-18
claude-sonnet-5
Anthropic
$0.270
$0.270
$4.91
2026-09-12
gemini-3-1-pro
Google
$0.307
$0.307
$5.58
2026-09-18
gpt-5-6-terra
OpenAI
$0.307
$0.307
$5.58
2026-09-12
gpt-5-6-sol
OpenAI
$0.540
$0.540
$9.81
2026-09-12
claude-opus-4-8
Anthropic
$0.675
$0.675
$12.27
2026-09-12
claude-opus-5
Anthropic
$0.675
$0.675
$12.27
2026-09-12
claude-fable-5
Anthropic
$1.349
$1.349
$24.54
2026-09-18
claude-fable-5-1
Anthropic
$1.349
$1.349
$24.54
2026-09-18
gpt-6-astra
OpenAI
$1.349
$1.349
$24.54
2026-09-17
Read this against the numbers above
Every row below prices the FAST TIER's own 55-call run (r001-appraisal-extract) -- the deliberating tier's own token counts are on Cost.cost_by_model[1] and are not separately projected here, and its output ran 70 pct higher (572.15 vs 335.80 tokens/call), the largest tier gap measured in this kit series so far.
Neither tier's registered run left anything to disable -- src/adapters/__init__.py's thinking parameter is only sent when a caller passes one, and this kit's own harness never does (see LLM.settings) -- so unlike several sibling kits, there is no reasoning-on/reasoning-off discrepancy to caveat here.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Five modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Three of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/segment.pysegment
cut the report into addressable sections, pure code
src/segment.py
# Cut a document into addressable sections. Pure code — no model, no network.
def sections(text):
def locate(text, value):
def span_label(secs, start):
src/select.pyselect — a swap seam
pick which sections carry each field, pure code — three sections for the extraordinary-assumption fields, since it can be stated in any of them
You change it to: map fields to your own report's headings; unmatched falls back to the whole document
src/select.py
# Pick which sections plausibly carry each field. Pure code -- the last deterministic step
SECTION_HINTS = {
def for_field(secs, field):
def plan(secs, fields):
src/prompt.pyprompt
assemble one call for all ten fields
src/prompt.py
# Assemble the extraction prompt. One prompt per report, all ten fields in it.
SYSTEM = (
def field_schema(fields):
def build(doc_text, secs, fields, selector):
def parse(raw, fields):
src/extract.pyextract
the AI layer, one provider one key — plus the pure-code computation downstream: the review-routing flag
src/extract.py
# Extract one report's fields: segment, select, prompt, one model call, then a pure-code
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FIELDS = os.path.join(HERE, "data", "fields.json")
CORPUS = os.path.join(HERE, "data", "corpus")
MAX_TOKENS = 4000
def load_fields():
def load_doc(stmt_id):
def documents():
def compute(values):
def extract(cfg, doc_text, fields, complete=None, thinking=None):
evals/judge.pyjudge
score field accuracy and the flag's recall/precision separately, pure code
evals/judge.py
# Score an extraction run. PURE CODE -- gold is exact and the answer is one value per cell, so
def norm(v):
def _num(v):
def equal(field, got, want):
def score_cell(field, got, want, stated):
def score(fields, records, golds):
def score_flags(flags, golds):
Start hereThe shortest path into it
src/segment.pycut the report into addressable sections, pure code
src/select.pypick which sections carry each field, pure code — three sections for the extraordinary-assumption fields, since it can be stated in any of them A swap seam.
src/prompt.pyassemble one call for all ten fields
src/extract.pythe AI layer, one provider one key — plus the pure-code computation downstream: the review-routing flag
evals/judge.pyscore field accuracy and the flag's recall/precision separately, pure code
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 774 input and 335 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's reports are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any report text. In a real deployment a report arrives from an appraiser or an AMC -- exactly the kind of externally-authored input this kit's architecture reads verbatim and trusts, with no verification step before the report's text reaches the model. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it -- and three of four boundaries hold
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every report is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's reports arrive from an appraiser or an AMC, and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it. Confirmed by reading the code, not by a run, on 2026-08-21 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does an extraction ever write anywhere or trigger a downstream action?
An extraction could plausibly write a USPAP compliance judgment, a value acceptance or an approval.
No code path does. src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/extract handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Is the review-routing rule something a prompt or a reply can move?
A crafted report could plausibly talk the routing rule itself into skipping a report, or fabricate a threshold that does not exist.
No. src/extract.py::compute() has no numeric constant at all -- it checks only whether the model's own extraordinary_assumption_present equals 'yes', read once at import time as pure code; nothing the model returns is consulted about how the rule itself works.
Could a crafted report talk the model into a fabricated extraordinary_assumption_present='no' with a real assumption stated in its text?
A report written to bury or disguise a real extraordinary assumption -- e.g. worded to read as a routine limiting condition rather than an extraordinary one -- could plausibly move extraordinary_assumption_present to 'no' with the assumption still present in the text.
Unmeasured -- no attack has been tried. extraordinary_assumption_present is an enum with no span at all (see Eval.scores.non_spannable_fields), unlike the copied-verbatim fields src/segment.py::locate() can check against the report text -- nothing stops the model from asserting a 'no' the report's own text does not support; see could_not_verify.
The first three boundaries hold, confirmed by reading the code, not by an attack trial. The fourth is the one this run has not tested: extraordinary_assumption_present is the field the whole guardrail depends on, and it is also the one field this kit's own span mechanism cannot check, because it is an enum, not a copied value -- no crafted input has ever been run against it.
The result0 attack trials, and three of four boundaries checked here hold in code: no write path exists, the routing rule has no threshold a prompt or a reply can move, and a misconfigured key cannot leak into a UI error. The fourth -- whether a crafted report could talk the model into a fabricated extraordinary_assumption_present='no' -- is unmeasured.
1externally-authored field a live deployment would carry (the report text itself, written by an appraiser or an AMC) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no report's text was authored by an outside party. A real deployment's reports arrive from an appraiser or an AMC -- exactly the kind of externally-supplied text this kit's architecture reads verbatim and trusts, with no verification step. Whether a report worded to disguise a real extraordinary assumption as a routine limiting condition could talk the model into a false extraordinary_assumption_present='no' is unmeasured for this kit.
Read this twice
The review flag is exactly as good as extraordinary_assumption_present. src/extract.py::compute() never re-derives extraordinary_assumption_present from the report text -- it trusts the model's own answer completely and routes on it alone, with no threshold to also get wrong. On this run both tiers scored a clean 1.0 (100 pct) recall and precision on the flag and zero errors on that field across 55 reports -- a genuinely clean run, and also a small one for a rule this consequential (see Business.not_good_enough). A future version should recompute extraordinary_assumption_present independently before trusting a live flag, and a red-team pass targeting exactly that field is the natural next measurement.
HonestyWhat this does not prove
Whether a report worded to disguise a real extraordinary assumption as a routine limiting condition could talk the model into a false extraordinary_assumption_present='no' -- no red-team run exists for this kit.
Whether the live app's own /api/extract behaves identically to the registered run under an adversarial report -- both use the same src/extract.py::extract(), but neither has been tested against one.
Whether a code-level consistency check on extraordinary_assumption_present (re-deriving it from the report text by pure code) would catch a model's fabricated 'no' in practice -- none has been built or exercised.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
An extraordinary assumption, stated anywhere in the report, routes it for review -- always, no dollar threshold. Computed in pure code from the model's own extracted extraordinary_assumption_present field, never asked of the model directly and never overridden by anything it returns. Unlike this kit's siblings (asset-verify, income-verify), which flag only above a $1,000 materiality threshold, this kit flags every stated assumption: its own facet sheet keeps USPAP compliance judgment and value acceptance with the certified review appraiser, never automated, so there is no dollar figure below which an assumption is judged immaterial here.
src/extract.py::compute() -- present == 'yes' is the entire rule, a module-level function read once at import time. Unlike a prompt instruction the model could ignore, this cannot be talked out of firing by anything in the reply text; it only depends on the model's own extraordinary_assumption_present value being right in the first place (see is_not and could_not_verify).
EvidenceDoes it hold?
What
Measured
The review flag never misses a report that should have been flagged
29 of 29 reports that should have been flagged were flagged, on both tiers (100 pct / 1.0 recall) -- see Eval.scores.
The review flag never fires on a report that should not be flagged
0 false positives among the 26 reports that should not have been flagged, on both tiers (100 pct / 1.0 precision).
There is no threshold a prompt or a reply can move -- every stated assumption routes, always
compute() has no numeric constant to move at all -- confirmed by reading the function; it checks only whether extraordinary_assumption_present == 'yes'.
No code path writes an approval, denial or value acceptance based on the flag
src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON extraordinary_assumption_present ITSELF. The flag trusts the model's own answer for that field completely -- nothing re-derives it from the report text before computing the flag. This run found zero disagreements on that field across 55 reports on both tiers (see Eval.scores), which is evidence the prompt's whole-report-read rule held up on this corpus, not evidence the guardrail itself checked and cleared anything -- it never looks at the report text at all.
It does not guarantee USPAP compliance or a value opinion -- the flag is a routing signal for a certified review appraiser, never itself a decision. See UI.shots and Data.breaks_on.
It has not been attacked. Whether a report worded to make a real extraordinary assumption read as routine could evade the flag is unmeasured -- see Security.could_not_verify.
WatchedWhat is watched, and why that one
2runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 14 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
8 measured by the latest run6 need the model half
Metric
Owner
Role
Why this one
field-exact-match-with-normalisation
Per-field exact match, light normalisation
alarm
extraction_accuracy specifically on extraordinary_assumption_text, since that is the one field with a real, repeatable gap -- see taxonomy; the 524-cell scored rate, since a drop anywhere would mean something in the prompt or the corpus changed; stated vs scored cells -- 524 of 550 possible cells were stated; the rest are correct nulls on no-EA reports, not misses — alarm on Any drop in extraction_accuracy below the measured 98.09 pct on either tier, or any change on extraordinary_assumption_present specifically (currently 55/55 on both tiers) -- that field is the one the review flag depends on entirely.
flag-confusion-matrix
Review-flag confusion matrix
alarm
review-flag recall specifically -- 29 reports should be flagged in this corpus, and missing one routes an unflagged extraordinary assumption past the certified review appraiser; false positives among the 26 reports that should NOT be flagged, since a false positive costs a reviewer's time rather than a missed risk; whether a field-level miss on extraordinary_assumption_text (see the other grader) ever changes the flag -- it has not, on either tier, this run, because the flag depends only on extraordinary_assumption_present, which had zero misses — alarm on Any review-flag false negative, on any run -- a missed flag is the safety-critical failure this kit exists to catch, and recall was 1.0 (100 pct) on both tiers this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
55
different corpus — nothing is comparable
corpus.bytes
38,168
reports edited — the count held, the bytes did not
split.count
504
the sections count moved — a different set was scored
split.size_p50
56
the median size of one section moved
split.size_p95
180
the 95th-percentile size of one section moved
dataset.rows
55
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.002
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (documents 55, extraction_cells 524, failures 0, refusal_cells 26, thinking True) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Extraction accuracy
exact match at 98.09 pct on both tiers
524 scored cells
two independent tiers (r001-appraisal-extract, r002-appraisal-extract) reproduced the identical figure to the digit -- the exact-match band MONITORING's own rule calls for.
Refusal accuracy
exact match at 100 pct on both tiers
26 refusal cells (correct nulls on reports with no stated extraordinary assumption)
two independent tiers reproduced the identical figure to the digit.
Hallucinations
exact match at 0 on both tiers
524 scored cells
two independent tiers reproduced zero to the digit.
Span rate
exact match at 86.71 pct on both tiers
spannable extracted values (non-enum fields with a non-null answer)
two independent tiers reproduced the identical figure to the digit.
Latency
2,366ms / 5,872ms p50/p95 on the fast tier, 4,823ms / 34,729ms p50/p95 on the deliberating tier -- the largest tier gap measured in this kit series so far, including one call at 154 seconds
55 calls per tier
measured directly on both tiers.
Token totals
42,604 input tokens on both tiers (identical prompt); 18,469 output on the fast tier vs 31,468 on the deliberating tier, about 70 pct more
55 calls per tier
measured directly on both tiers.
HistoryRun history
2 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-appraisal-extract 2026-08-21
r002-appraisal-extract 2026-08-21
extraction accuracy
0.9809
0.9809
refusal accuracy
1.000
1.000
invented values
0
0
values with a span
0.8671
0.8671
input tokens, whole run
42604
42604
model latency p50 ms
2366.00
4823.00
model latency p95 ms
5872.00
34729.00
output tokens, whole run
18469
31468
not a time series No two of these 2 runs measured the same system — they differ on the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 2 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
which tier is called (fast vs deliberating)
output tokens (+70 pct) and latency (roughly 2x p50, nearly 6x p95) move together; extraction accuracy, refusal accuracy, hallucinations, span rate and the review flag's recall/precision do not move at all
extraordinary_assumption_present resolves to a correct 'no' and extraordinary_assumption_text resolves to a correct null together, and computed needs_review resolves to false rather than being left unset
reasoning
src/extract.py::compute() returns None only when present itself was never extracted, a third state distinct from a confident 'no' -- confirmed by reading the function; this run's corpus includes reports with no extraordinary assumption by construction (NO_EA_FRACTION=0.40 in tools/build_corpus.py) but no committed run isolates their scores separately.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Nothing fires yet. Not one of the 6 bands on this kit names a consequence, so there is no row to print. Which is the honest answer to the heading: a band that nothing is wired to is a number somebody has to notice, and this kit has only those. An empty table would have said the same thing while looking like a rendering failure.
NextThe three you would add first
Re-derive extraordinary_assumption_present from the report text by pure code, and compare it to the model's own answer before trusting the flagthe flag's only real dependency is a field this guardrail currently takes on faith -- src/segment.py already does word-boundary text matching for spannable fields, though extraordinary_assumption_present is itself non-spannable (see Eval.scores.non_spannable_fields) so a comparable check would need its own approach.
Cross-reference the report against its own addenda before trusting that no extraordinary assumption was missedthis corpus never places one only in an addendum the main body references but does not repeat -- see Data.breaks_on -- and a real report's addenda are not read here at all.
Run a red-team pass against extraordinary_assumption_present specificallyit is the one field the whole guardrail depends on, and it is the one field this run has not attacked -- see Security.could_not_verify.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-check the flag's evidence on any change to the present == 'yes' rule in src/extract.py::compute(), or to the extraordinary-assumption instruction in src/prompt.py -- either changes what the flag is computed FROM. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
What this cannot tell you
Whether the flag would still hold 1.0 recall and precision on a larger or harder set of embedded assumptions -- this run's 11 embedded cases (of 29 stated assumptions) all resolved correctly on both tiers, but 11 cases is not enough to rule out a harder placement this corpus did not think to plant.
Whether a report worded to make a real extraordinary assumption read as routine, or a prompt-injection attempt inside a report section, could suppress extraordinary_assumption_present or needs_review -- no red-team run exists for this kit, see Security.could_not_verify.
Why the deliberating tier's output tokens and tail latency were so much higher on this corpus specifically -- see Eval.could_not_verify; this is provider-default behaviour, not a setting this kit's own harness controls.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is empty, with a comment explaining that the emptiness is load-bearing. The whole extraction decision is three files: src/segment.py, src/select.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
55 appraisal reports, generated from a fixed seed, never fetched. Gold is read back off the same generated report text, never carried over from a separate target -- evals/check_labels.py asserts stated/value and present/text consistency on every row before any run is allowed to spend; see data/SOURCES.md.
segmentation and selection
src/segment.py, src/select.py
text splitters / retrievers
a heading-based cut (src/segment.py) and a fixed field-to-heading map (SECTION_HINTS in src/select.py) -- no embeddings, no index, no ranking; the extraordinary-assumption field pair is the one sent three sections rather than one, since the assumption can be stated in any of them. A field absent from the map gets the whole document rather than nothing.
prompt assembly
src/prompt.py
prompt templates
the ten-field schema and the whole-report-read rule are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. This kit's own MAX_TOKENS=4000 (carried over from sibling kit docs-extract's own experience on a similarly-shaped record) is a plain constant, not a client-library setting.
evaluation
evals/judge.py
eval harnesses
per-field exact match plus a separately-scored review-flag confusion matrix is a loop and a handful of counters, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per report -- segment, select, prompt, call, parse, compute -- with no branching and no state carried between reports. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A different field schema (data/fields.json) or a different routing rule needs its own gold set and its own evals/check_labels.py pass -- a framework's own schema layer would not remove that work, only relocate it.
SECTION_HINTS is a five-minute edit for a new report layout because it is a plain dict, not a configured retriever -- a framework's own chunking/retrieval abstraction would need its own re-tuning pass instead, with its own failure modes to learn.
Swapping providers is one function and one entry in PROVIDERS (src/adapters/__init__.py) -- a framework's own model abstraction would add a dependency and a version to track for the same one-line change this file already gives away free.
What we could NOT verify
Whether a framework's own retrieval or agent abstraction would resolve anything this kit does not already handle was not tested -- this run found zero errors on extraordinary_assumption_present on either tier, and the one real gap (extraordinary_assumption_text narrowing) is a text-boundary choice, not a retrieval failure (see Eval.taxonomy), so there is no observed weakness here for a framework to address.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-appraisal-extract on the fast tier, 2026-08-21. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
2,366 ms
2,366ms / 5,872ms p50/p95 on the fast tier, 4,823ms / 34,729ms p50/p95 on the deliberating tier -- the largest tier gap measured in this kit series so far, including one call at 154 seconds
—
Model, p95
5,872 ms
2,366ms / 5,872ms p50/p95 on the fast tier, 4,823ms / 34,729ms p50/p95 on the deliberating tier -- the largest tier gap measured in this kit series so far, including one call at 154 seconds
—
Input tokens
42,604
42,604 input tokens on both tiers (identical prompt); 18,469 output on the fast tier vs 31,468 on the deliberating tier, about 70 pct more
—
Output tokens
18,469
42,604 input tokens on both tiers (identical prompt); 18,469 output on the fast tier vs 31,468 on the deliberating tier, about 70 pct more
—
No movement column. Not one of the 1 earlier run on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-appraisal-extract2,366 ms
r002-appraisal-extract4,823 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-21, across 2 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
appraisal reports
data/corpus/*.txt -- 55 reports, generated once from a fixed seed (SEED = 20260821) by tools/build_corpus.py
read whole by src/segment.py and src/select.py; never modified after generation
gold labels
data/gold.jsonl -- 55 rows, read back off the same generated report text (never a target that seeded them), checked against the document text by tools/build_corpus.py's own _verify() pass
never -- evals/judge.py is pure code, no model, no key
the field schema
data/fields.json -- the ten-field record src/prompt.py assembles the user message from
read by src/prompt.py and src/extract.py only
the key
.env -- never committed (see .gitignore)
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
model
one call per REPORT, carrying the selected sections plus the ten-field schema, behind src/adapters/__init__.py -- OpenAI-compatible wire format over raw HTTP, so a forker runs this on whichever key they hold. MAX_TOKENS is fixed at 4000 (src/extract.py), set on sibling kit docs-extract's own experience with a similarly-shaped JSON record rather than a ceiling this kit has ever hit -- 0 of 55 calls truncated on either tier.
2,366ms p50 / 5,872ms p95 on the fast tier, 4,823ms p50 / 34,729ms p95 on the deliberating tier -- the largest tier gap measured in this kit series so far, including one call at 154 seconds (the fast-tier and deliberating-tier runs, 2026-08-21 -- see Cost.cost_by_model)
one call per report, no concurrency and nothing shared between calls -- see Architecture.breaks_at_scale. A stack of thousands needs batching and a rate-limit strategy this kit does not have; MAX_CALLS_PER_DAY in src/budget.py caps the shared key across every kit on this machine, not this kit's own throughput.
point src/adapters/__init__.py at a different provider or model and every published accuracy/cost/latency figure is void until re-run -- this run's numbers are this model's numbers, not a property of the prompt.
corpus refresh
tools/build_corpus.py regenerates the whole corpus -- 55 appraisal reports across a fixed roster of appraisers, properties and extraordinary-assumption placements -- byte-identically from a fixed seed (SEED = 20260821) every time it is run. There is no incremental refresh; gold is derived from the same generated report text (never a target that seeded it) and the script's own _verify() pass checks every report's stated dollar figure and extraordinary-assumption text appear verbatim in the document text before anything is scored.
0.041s wall time to regenerate all 55 reports and their gold labels (measured directly, 2026-08-21 (python3 tools/build_corpus.py, timed) -- see Data.index for the separate segmentation figure, a different step)
a real deployment's report layout, appraiser roster and extraordinary-assumption vocabulary do not come from a fixed seed and grow without bound -- how a genuinely varied set of real-world appraisal-report templates would change what src/segment.py (heading-based) and src/select.py (SECTION_HINTS) can resolve was not measured; see Architecture.breaks_at_scale and Data.breaks_on.
point tools/build_corpus.py at your own report layout and appraiser roster, and every published accuracy figure is void -- they are this corpus's own planted ambiguity, not a property of the model.
labels
evals/check_labels.py asserts gold/document consistency (one gold row per document and vice versa, every enum value allowed by the schema, and 'stated' agreeing with whether extraordinary_assumption_text is null, and extraordinary_assumption_present agreeing with whether the text is null) before evals/run.py is allowed to spend anything. Scoring (evals/judge.py) is per-(report,field) exact match with light normalisation, plus a separately-scored review-flag confusion matrix computed from the model's own extracted fields by pure code (src/extract.py::compute), never from a second model call.
524 of 550 possible (report,field) cells actually scored (55 reports x 10 fields; a null extraordinary_assumption_text on a no-EA report is a correct abstention, not a cell to score) plus 29 reports carrying a true review-flag verdict (needs_review=true), scored separately (lenses.Eval.dataset, the fast-tier and deliberating-tier runs, 2026-08-21)
the gold set stops at 55 reports and covers one planted ambiguity (an extraordinary assumption embedded in prose vs. stated under its own heading) at a fixed 40 pct rate -- a real portfolio's mix of placements is unmeasured, and an addendum the main body references but does not repeat is never tested; see Data.breaks_on.
a different field schema (data/fields.json) or a different routing rule needs its own gold set and its own evals/check_labels.py pass before any published figure can be trusted again.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
extraordinary_assumption_present answered 'yes' on a report with no dedicated 'Extraordinary Assumptions' heading at all
the model read the whole report rather than checking only a labelled section -- this is the corpus's planted ambiguity (AMBIGUOUS_FRACTION=0.40 of stated assumptions in tools/build_corpus.py, embedded in Scope of Work or Comments prose instead), and on this run both tiers got every one of the 11 embedded cases right (see Eval.taxonomy)
read the report's Scope of Work and Comments sections before assuming a 'yes' with no visible Extraordinary Assumptions section is a mistake -- the prompt's own rule (see README) is to check the whole report, and this run found zero cases where either tier missed an embedded assumption (the fast-tier and deliberating-tier result files, both runs, 2026-08-21 -- see Eval.taxonomy (the text-boundary-narrowed entry) and Eval.baseline for the same pattern's effect on the free heading-only floor)
No machine symptom — this failure leaves no trace in any output.
no path in src/extract.py::compute() or src/app.py writes a USPAP compliance judgment, a value acceptance or an approval of any kind -- needs_review is returned to the caller as a field on the response record and nothing downstream of this kit acts on it. A missed flag on a real deployment would show up only in whatever collateral-review system consumes this kit's output, which this kit does not have and does not simulate -- so there is no committed artifact naming that failure, on purpose: it is out of this kit's boundary, not unmeasured.
Whether a 55-report, single-seed run's zero-error result on extraordinary_assumption_present generalises to a harder embedded placement this corpus did not plant, or to a larger sample size, was not tested -- see Eval.could_not_verify. Concurrency (every run in this series is one call at a time, sequential), whether the no-threshold routing rule matches any real review-appraisal program's own materiality policy, and how either tier performs on a real, messy appraisal-report layout (multi-page PDFs with addenda, form-based URAR layouts, scanned images with no OCR step) are all unmeasured -- see Eval.could_not_verify and Data.breaks_on for the full list.
The corpus licence, from the Data lens: MIT — this repository's own licence. Every property, appraiser and dollar figure is invented; a real appraisal report cannot be published either — it carries a real address and a real licensed appraiser's identity tied to a real client's loan file. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one linePer-field exact match, light normalisation
Does the model's value for each field match gold, after trimming whitespace/punctuation and treating numbers within half a cent as equal?
$0.00per 1,000 reports
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
doc
AR-0003
field
extraordinary_assumption_present
extraordinary assumption present model
yes
gold extraordinary assumption present
yes
needs review
yes
found in section
Scope of Work (no dedicated heading present)
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
AR-0003 states an extraordinary assumption about the subject's septic system, embedded mid-sentence in the Scope of Work section with no dedicated 'Extraordinary Assumptions' heading anywhere in the report. Gold: extraordinary_assumption_present=yes. Both tiers answered yes and the run's own pure code then computed needs_review=true, correctly routing it for review despite there being no labelled section to key off.
Review-flag confusion matrix
correct
AR-0003: gold needs_review=true (the septic-system extraordinary assumption is stated in the report, embedded in Scope of Work with no dedicated heading). Both models' own extracted extraordinary_assumption_present='yes' fed src/extract.py::compute(), which correctly computed needs_review=true on both tiers -- routed for review as it should be.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 98.1%
the deliberating tier
scored 98.1%
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold, read back off the same generated report text (never a target that seeded it); the internal consistency check (_verify()) confirms every stated dollar figure and extraordinary-assumption text appears verbatim in the document text before anything is scored.
These rates are UNKNOWN, on purpose
This grader IS the reference standard for field values -- it cannot be scored against itself. What can go wrong is the corpus's own generation logic, which tools/build_corpus.py's own _verify() pass checks by confirming every report's stated figures and extraordinary-assumption text appear verbatim in the document text.
Watch these
extraction_accuracy specifically on extraordinary_assumption_text, since that is the one field with a real, repeatable gap -- see taxonomy
the 524-cell scored rate, since a drop anywhere would mean something in the prompt or the corpus changed
stated vs scored cells -- 524 of 550 possible cells were stated; the rest are correct nulls on no-EA reports, not misses
Alarm on
Any drop in extraction_accuracy below the measured 98.09 pct on either tier, or any change on extraordinary_assumption_present specifically (currently 55/55 on both tiers) -- that field is the one the review flag depends on entirely.
How tight can the band be? There is no tolerance band on the field grade itself for eight of the ten fields -- exact match after trimming whitespace/punctuation and treating numbers within half a cent as equal. extraordinary_assumption_text is scored the same exact-match way and the 10-of-29 narrowed-span cases score as 'wrong', not partial credit for a boundary choice -- see taxonomy.
Cadence: Re-score on any change to src/prompt.py, data/fields.json or tools/build_corpus.py -- the first two change what is asked, the third changes what is asked about. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
The decisionWhen to reach for it
Use it
Gold is derived from the same generated report text, never from a separate target -- true of every kit corpus, never true of a real collateral review team's own report history.
Do not use it
The true field values are not known in advance -- the normal state of a real appraisal review, and the reason this corpus is generated rather than captured.
Catch buried assumptions in a mortgage lender's appraisal reports
PresenterOpens the private repo. Visible to admins only.
In one lineReview-flag confusion matrix
Does the run's own pure-code needs_review (computed from what the model extracted for extraordinary_assumption_present) match the same computation run over gold's own true values?
$0.00per 1,000 reports
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model. Computed from the run's own pure-code needs_review (src/extract.py::compute), never from a second model call -- the model never sees or states the flag directly.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
doc
AR-0003
field
extraordinary_assumption_present
extraordinary assumption present model
yes
gold extraordinary assumption present
yes
needs review
yes
found in section
Scope of Work (no dedicated heading present)
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
AR-0003 states an extraordinary assumption about the subject's septic system, embedded mid-sentence in the Scope of Work section with no dedicated 'Extraordinary Assumptions' heading anywhere in the report. Gold: extraordinary_assumption_present=yes. Both tiers answered yes and the run's own pure code then computed needs_review=true, correctly routing it for review despite there being no labelled section to key off.
Review-flag confusion matrix
correct
AR-0003: gold needs_review=true (the septic-system extraordinary assumption is stated in the report, embedded in Scope of Work with no dedicated heading). Both models' own extracted extraordinary_assumption_present='yes' fed src/extract.py::compute(), which correctly computed needs_review=true on both tiers -- routed for review as it should be.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 100.0%
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold needs_review, computed the identical way src/extract.py::compute() computes it -- from the same generated report text, using present == 'yes' as the entire rule (no threshold).
These rates are UNKNOWN, on purpose
This grader is scored against a mechanically-derived gold value (src/extract.py::compute(), the same pure-code path both the run and the gold set use), not against a second grader's judgement -- there is nothing here to publish an agreement rate against. What can go wrong is the corpus's own placement of extraordinary assumptions, which tools/build_corpus.py's own verification pass checks against the same generated report text.
Watch these
review-flag recall specifically -- 29 reports should be flagged in this corpus, and missing one routes an unflagged extraordinary assumption past the certified review appraiser
false positives among the 26 reports that should NOT be flagged, since a false positive costs a reviewer's time rather than a missed risk
whether a field-level miss on extraordinary_assumption_text (see the other grader) ever changes the flag -- it has not, on either tier, this run, because the flag depends only on extraordinary_assumption_present, which had zero misses
Alarm on
Any review-flag false negative, on any run -- a missed flag is the safety-critical failure this kit exists to catch, and recall was 1.0 (100 pct) on both tiers this run.
How tight can the band be? There is no tolerance band -- the flag is a binary confusion matrix (flagged/not flagged against gold), never a continuous quantity to round.
Cadence: Re-score on any change to the present == 'yes' rule in src/extract.py::compute() or to the extraordinary-assumption fields in tools/build_corpus.py or data/fields.json. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
The decisionWhen to reach for it
Use it
Gold's needs_review is computed the identical way from the same generated report text -- true of every kit corpus, never true of a real review appraiser's own judgment call on materiality.
Do not use it
The true routing determination is not known in advance, and this kit's own no-threshold rule (every stated assumption routes, unlike its $1,000-threshold siblings) has not been checked against a real program's own review policy -- see Eval.could_not_verify.
A living map of modern AI — kept current every morning