Home › Use Cases › Cross-check two independent quantity takeoffs of one work package, element by element
Use caseUC0286
🧪 Use-case kit · runnable

Cross-check two independent quantity takeoffs of one work package, element by element

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Two estimators measure the same drawings and produce two sheets that share no wording, no order, no line breakdown and sometimes no unit. Somebody then has to say where the two disagree, and the hard part is not the arithmetic — it is deciding which line on one sheet is the same physical element as which line on the other. Do it by eye on a nine-line package and it takes twenty minutes; do it wrong and you either send an estimator after a gap that does not exist or let a real one into a bid. Nothing is replaced. It produces the report an estimator confirms, and it makes no decision anybody currently makes.

Audience

A preconstruction lead or chief estimator cross-checking a subcontractor's measure against their own, or one takeoff against its own earlier version after a drawing revision. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual cross-check package

The corpus is 62 cross-check package, 0.20 MB (txt 62). Because a cross-check has exactly one reading problem in it and a corpus can hand that problem away four different ways without noticing. It did, three times, and each was measured and closed before a call was bought: both sheets printed the same classification code (a regex grouping on it solved all 62 packages); both sheets carried the same number for every element (exact-quantity matching solved them); and every element had a distinct order of magnitude (nearest-quantity solved 58 of 62). The fourth was the decoy, which defeated neither free signal in its first form. What is left is a corpus where the free floor reads 54 of 62 — and that is still not hard enough, because the paid arm reads 62.

The corpus

  • The 62 cross-check packagegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states what is invented (all of it), what the generator costs the measurement, and the four free answers the corpus was hardened to close.

Swap this folder for your own material and the kit is pointed at your cross-check package. That is the whole change — there is no database to migrate.

One cross-check package, as the model receives itTCP-0001.txt · 1 of 62
TAKEOFF CROSS-CHECK PACKAGE                                         TCP-0001
Prepared 2026-09-02 under TCC-2026 | two independent takeoffs of one work package

PACKAGE FACTS
  Project                      Tarnholm Transit Depot
  Work package                 Roof and weatherproofing
  Package reference            TCP-0001
  Takeoff A measured by        Marlow Quantities
  Takeoff B measured by        the trade contractor's own measure

UNIT BASIS
  Area                         SF is the base. 1 SY = 9 SF. 1 SQ = 100 SF.
  Length                       LF is the base. No other length unit is used.
  Volume                       CF is the base. 1 CY = 27 CF.
  Weight                       LB is the base. 1 TON = 2000 LB.
  Count                        EA is the base.
  Not convertible              a quantity on one basis against a quantity on another - a wall in LF
                               against the same wall in SF - has no factor on this package.

CROSS-CHECK BASIS
  Material divergence          more than 2.0 pct of the larger side
  Immaterial below             25.000 SF area | 10.000 LF length | 5.000 CF volume
                               50.000 LB weight | 1.000 EA count
  Both must be cleared         a difference under either half is not material

TAKEOFF A LINES        measured against the coded specification
  id     code      description                                                                    quantity  unit
  A-01   04 22 00  Exterior wall, concrete masonry unit, 8 in                                    4,139.214  SF
  A-02   05 12 00  Structural steel framing, wide flange members                                71,832.864  LB
  A-03   05 31 00  Steel roof deck, 1.5 in wide rib                                              4,105.276  SF

Abridged — the file continues.

The outcomeWhat a good result looks like

One report per work package: every line on both sheets in exactly one element group, a finding per element, both quantities on one basis, the difference and the percentage, and the package's status. The estimator adjudicates every divergence; the report names none of them right.

And when it cannot

The report says AGREES on a package that does not. That is the only outcome here that lets a wrong quantity into a bid with nobody left to catch it, and it is why AGREES is LAST in the status vocabulary rather than first. Measured at 0 of 62 on the paid arm, 0 of 62 on the free rules floor, and 36 of 62 on the null floor — which is what an arm that reads nothing does.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • a subcontractor's measure against your own, on one work package — this kit, whole
    it is the shape the kit was built for: two sheets, no shared template, neither authoritative
  • one takeoff against its own earlier version after a drawing revision — this kit, with the register set to revision_gap
    TX-1 says the cross-check is not yet meaningful and the element findings are reported underneath it rather than acted on
  • deciding which of two quantities to carry into a bid — a person. The report names the element, both numbers and the gap; somebody who knows the job decides which is right.
    the answer contract has no field that could express a preference and src/prompt.py fails at import if one is added
  • pricing the difference once it is found — a costing tool, downstream of this one
    no rate, no extension and no money appears anywhere in the contract or the report

And where nothing here is good enough:

  • a two-hundred-line package — nothing measured. This is unmeasured ground, not a refusal.
    these packages carry 6 to 11 lines a side and nothing here blocks or shards. The pairing is quadratic and it is the part that would move first

At a glanceHow the whole thing runs

62%package all correct (rechecked)
45,055 msp50, end to end
$23.81per 1,000 cross-check package · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own two-sheet packages in the same six-panel layout, or rewrite tools/build_corpus.py's render() for yours and keep the answer key derived. The answer key is the boundary. Corpus lens →
When is this the wrong choice?Avoid: Paying for the arithmetic. Every quantity, conversion, difference and percentage here is integer code over numbers already printed, and the free floor gets all of it for $0.00 — what the call buys is the pairing, and only where the two sheets share no vocabulary or carry a near twin. That is the case against the best-fitting scenario (“a subcontractor's measure against your own, on one work package”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a scanned or photographed takeoff — everything here is machine-readable text. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second run of the same corpus. Every figure is one run. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-takeoff-crosscheck. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run. Both free floors, the rulebook, the register, the unit arithmetic, the board and every committed run need no key and no network. Only evals/run.py without --floor/--stub, evals/injection.py and the board's one model button reach a provider.

A living map of modern AI — kept current every morning