The business caseThe problem this solves
Two estimators price the same job separately and the two documents share no line order, no wording and no granularity: one writes 'Refinish and blend front bumper cover, 3.3 hrs' where the other writes 'Paint bumper cover, front, 2.5' and 'Blend into adjacent panel, 0.8'. Deciding those are the same work is reading. Everything after that decision -- each side's totals, the quantity, the hours, the money, which of four kinds the group is, and the net -- is arithmetic. So somebody sits with two printed documents, a ruler and a calculator, and builds the comparison by hand. Building the line-by-line comparison of two estimates by hand -- reading which line answers which, then adding up each side, the deltas, the buckets and the net. It replaces the building; it does not replace the confirming, and there is no endpoint that approves, adjusts or settles anything (CMP-1.11).
Audience
Whoever has to sign off the difference between two estimates for one job -- the assessor or adjuster who confirms the comparison, and the person who answers for the money that moves because of it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual pairs of estimates
The corpus is 60 pairs of estimates, 0.24 MB (txt 60). The defect mix a comparison actually fails on, and no public dataset can provide it. 15 pairs carry a split or a merge -- the case a one-to-one matcher cannot express at ALL, which is what separates a real comparison from a lucky one. 4 renumber the B side's parts into another supplier's catalogue, so the two codes share no characters and the description is the only evidence left. 4 reword part-less lines out of recognition. 4 plant lookalike lines one position apart -- left headlamp against right -- which is the most common way a comparison goes quietly wrong, because crossing them makes two real findings vanish into one false agreement. 3 price the same hours at different labour rates. And 33 pairs are clean, because a corpus that is all traps measures a different job from the one the desk has: 480 of the 547 groups agree, which is the shape of the real problem -- finding the handful of lines that do not, without drowning them in a hundred that do.
The corpus
- The 60 pairs of estimatesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your pairs of estimates. That is the whole change — there is no database to migrate.
JOB RECORD
Job reference JOB-4400
Item 2019 five-door hatchback, front nearside impact
Estimates 2, prepared independently. Neither is authoritative.
Amounts printed to two decimal places; compared as whole minor units.
==============================================================================================================================
ESTIMATE A
Prepared by Merrivale Assessing Services
Prepared on 3 January 2026
Labour rate 65.00 per hour
Lines 9
Ref Op Description Part no. Qty Hours Unit price Amount
------------------------------------------------------------------------------------------------------------------------------
A-01 REP Radiator support panel - repair and align - 0 2.8 0.00 182.00
A-02 RR Fog lamp right - replace 81210-0K040 1 0.4 126.00 152.00
A-03 RR Front wing right - replace 53811-0K180 1 2.2 276.00 419.00
A-04 RR Wheel arch liner front left - replace 53876-0K060 1 0.6 84.00 123.00
A-05 RR Air conditioning condenser - replace 88460-0K270 1 1.1 349.00 420.50
A-06 RR Bonnet - remove and replace 53301-0K120 1 1.8 612.00 729.00
A-07 RR Driver airbag module - replace 45130-0K190 1 1.2 745.00 823.00Abridged — the file continues.
The outcomeWhat a good result looks like
A comparison a person CONFIRMS instead of one they build: every line on both documents in exactly one group, each group's kind and its money, the three buckets the net decomposes into, and one sentence saying what the difference actually is -- with both sides of every group spelled out line by line so the reader can check it rather than believe it.
And when it cannot
A real difference reported as agreement. A comparison's whole output is the list of things that differ, so a difference that does not make the list is not a small error -- nobody looks at it again. This corpus puts 9,513.25 of genuine difference in 67 groups: the free rules floor misses 1 of them (26.00) and invents 74 that do not exist, carrying 7,776.20 of money that is not at stake -- more than the entire genuine total. The model missed 0 and invented 0 on this corpus, which is the finding and also the reason the finding is thin.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Two systems that always itemise the same way -- one estimating package on both sides, or a feed where the line breakdown is fixed by a schema — the free rules floor alone
Its entire loss on this corpus is granularity: 8 of its 26 misses are split_merge_missed and most of the other 17 are the same failure one step downstream. Where no estimate combines what the other itemises, part-number normalisation and a description threshold do the job for $0.00, and it already gets 99.7 pct of the money in real differences. - Two independently prepared estimates that split the job differently -- merges, splits, redrawn boundaries, parts renumbered into another supplier's catalogue — the model, rechecked
15 of 15 split/merge groups against the floor's 7 of 15, and 547 of 547 groups on both precision and recall against 93.4 pct and 86.9 pct. The floor draws 588 groups where the key has 547, and every extra one is a merge it could not express reported as an omission plus an addition. The recheck keeps the alignment and re-derives every figure, so a model's arithmetic mistake never reaches the comparison.
And where nothing here is good enough:
- Estimates that arrive as scanned pages, faxed PDFs or photographs of a printout — neither arm, until something else has read the page
This kit starts after somebody turned two estimates into line-item records. Reading a table off a page is a different problem with a different failure mode, and a kit that did both would report one number for two jobs.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own two estimates into data/corpus/ in the printed layout this kit uses -- job record, then each estimate with its own labour rate in its header and one row per line: ref, operation code, description, part number, quantity, hours, unit price, amount. ⚠︎ BOTH ESTIMATES REACH YOUR CONFIGURED PROVIDER VERBATIM, AND ONE OF THEM WAS WRITTEN BY SOMEBODY WITH MONEY RIDING ON THE COMPARISON. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per comparison for a job a regex, a set measure and an integer subtraction already do -- and inheriting a 34.5-second p50 to do it. That is the case against the best-fitting scenario (“Two systems that always itemise the same way -- one estimating package on both sides, or a feed where the line breakdown is fixed by a schema”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | REAL ESTIMATES. The documents are templated: the descriptions come from a pool of two or three wordings per component, so the surface variety of two real estimating systems -- abbreviations neither side expands, an operation one shop bills as sublet and the other in-house, a line that is three words and a price -- is not in it. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHERE THIS MODEL'S EDGE IS. Every failure bucket is empty and every graded field is 100 pct, so the run says what did not happen and cannot say what would. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-estimate-compare. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on a copy of the kit with no .env and no API key: python3 tools/build_corpus.py rebuilt the 60 pairs, the 1,091 line items and the 547-group key in 0.06 s, BYTE-IDENTICAL to the committed corpus (same md5 on gold.jsonl, estimates.json, jobs.json and corpus-stats.json; diff -rq over data/corpus reports no difference). python3 -m evals.check_labels re-graded that key with independent arithmetic in 0.05 s -- KEY CLEAN, 60 pairs, 547 groups. python3 -m evals.run --floor rules scored the free floor at 34 of 60 in 0.11 s. python3 -m src.app served the board with the model button disabled and saying why. requirements.txt names nothing, so there is no install step between a clone and all of that.



