Home › Use Cases › Check one quote against the pricing policy in force before it is signed
Use caseUC0358
🧪 Use-case kit · runnable

Check one quote against the pricing policy in force before it is signed

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A seller sends a quote to a deal desk and somebody has to say, before it is signed, whether it is inside the pricing policy. That is nine separate questions on one document — is any line under its floor price, is any discount over the tier's ceiling, does the approval on file actually cover the discount, are the payment terms and the term length ones this tier permits, does the auto-renewal state a notice period inside the window, does each ramp year rise by the minimum uplift, are the units given away inside the allowance, and is there a term on the restricted list that this discount does not permit. Seven of the nine are a lookup. Two are a reading: an approval record is sentences, and a restricted term is identified by what a clause DOES rather than by whether it uses the name. Reading every quote line against a published floor price and a published discount ceiling, every approval-record sentence against an approval matrix, and every special term against a restricted list and a threshold that moves with the discount — and writing down, for each of the nine, the numbered policy line it was checked against.

Audience

A deal desk or revenue-operations reviewer who sees quotes before they go out, and the sales operations person who has to say afterwards which quotes went out of policy and against which line. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual quotes awaiting signature

The corpus is 63 quotes awaiting signature, 0.20 MB (json 3 · jsonl 1 · md 2 · txt 63). Because the shape of a pricing-policy failure is not the arithmetic, it is the sentence. Seven of the nine checks are a comparison against a number printed on the quote itself. The two that are not — what an approval record actually establishes, and what a special term actually does — are the only place a model can earn its bill, and a corpus that did not build both of them in BOTH directions would measure a checker that flags everything or one that flags nothing.

The corpus

  • The 63 quotes awaiting signaturegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 63 quotes, data/pricebook.json and the whole answer key are written by tools/build_corpus.py. There is no real customer, no real contract, no real quote and no real price book anywhere in this kit, and no personal data either.

Swap this folder for your own material and the kit is pointed at your quotes awaiting signature. That is the whole change — there is no database to migrate.

One quotes awaiting signature, as the model receives itQP-0001.txt · 1 of 63
QUOTE - PRICING POLICY CHECK

QUOTE HEADER
  Quote number             QP-0001
  Opportunity              OPP-31000
  Customer                 Aldenmere Logistics Group
  Vendor                   Northlight Software
  Tier                     Standard
  Quote date               2026-08-03
  Subscription term        24 months
  Payment terms            Net 30
  Auto-renewal             Yes, 45 days notice
  Approval claimed         Sales manager
  Special terms declared   None
  Raised by                R. Okonjo, for Northlight Software

PRICING POLICY EXTRACT (QPOL-2026 r3, in force 2026-07-01 to 2026-12-31 - the Standard column)
  Maximum discount off list        20.0 pct
  Permitted payment terms          Net 30
  Permitted term lengths           12, 24 months
  Minimum annual uplift            5.0 pct
  Units at no charge, allowance    2.0 pct of the units charged for
  Restricted terms, threshold      10.0 pct off list
  Approval matrix                  to 10.0 pct none, over 10.0 sales manager, over 20.0 sales director, over 30.0 regional vice-president, over 40.0 chief financial officer
  Notice window for auto-renewal   30 to 90 days
   SKU       Product                         List unit      Floor unit
   PLT-CORE  Platform Core seat                $319.00         $239.25
   PLT-PLUS  Platform Plus seat                $425.00         $374.00
   ANL-INS   Insights Analytics seat           $288.00         $216.00
   WFL-ORCH  Workflow Orchestrator module      $456.00         $342.00
   SEC-GOV   Security Governance module        $334.00         $293.92
   INT-BUS   Integration Bus connector         $295.00         $221.25
   SUP-ELV   Elevated Support pack             $249.00         $186.75
   DAT-RET   Data Retention block              $312.00         $234.00

Abridged — the file continues.

The outcomeWhat a good result looks like

All nine checks answered on every quote, each carrying the NUMBERED POLICY LINE it was checked against and, where it is not met, one row copied verbatim out of the quote. Rechecked in code from the model's two readings the kit takes 566 of 567 findings, 567 of 567 numbered policy lines and 63 of 63 recommendations, and names the money below the floor to the cent on 63 of 63 quotes.

And when it cannot

⚠︎ THE QUOTE-LEVEL NUMBER IS 18 OF 63 AND IT IS NOT A TYPO. quote_all_correct demands all nine findings, both readings, the recommendation, the breached set, the money AND a citation on every finding that needs one and NOTHING quoted on the 525 that do not. The arm returned a row on 66 of the 525 checks that are MET or NOT-ENGAGED, where the answer contract and the quoting rule both say null. Every row it quoted was real — 104 returned, 0 unlocatable — so this is over-citing, never invention, and the pure-code station does not repair a citation on purpose. That one behaviour is the whole gap between 566 of 567 findings and 18 of 63 quotes.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your sellers code Approval claimed and Special terms declared honestly, and you only want the seven checks a lookup settles. — the free column floor alone — python3 -m evals.run --run-id b000-<slug>-rules --floor rules. It takes 549 of 567 findings and 42 of 63 quotes for $0.00.
    Every one of those seven is a comparison, a set membership test or two divisions in whole basis points. Nothing about them needs a model.
  • Your approval record is free text and your sellers' header codings are optimistic. — the paid call, and read the rechecked column. It takes 18 of the 18 checks whose answer is in prose against the column floor's 0, exact two-sided p = 0.0000076.
    That sub-population is the only place on this corpus where any arm separates from any other on evidence rather than on memorisation.
  • You need every finding to name the policy line and quote the row, because that is your audit standard. — the paid call for the status and the policy line — 566 of 567 and 567 of 567 rechecked — and a code post-pass that DELETES a citation on any finding that is met or not engaged.
    The arm's citations are all real (104 returned, 0 unlocatable). The failure is entirely that it quotes rows it should not, and that is a deletion rather than a repair.
  • You want a number for how much money is under the floor. — the rechecked column, which is 63 of 63 quotes to the cent — or pure code, which is the same answer for nothing.
    It is one subtraction and one multiplication per line and no reading enters it. The raw arm got 49 of 63 and overstated by $18803.30.

At a glanceHow the whole thing runs

99.8%check status correct rechecked pct
3,184 msp50, end to end

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported quotes and data/pricebook.json at your own price book, rewrite data/policy.md and data/policy.json as YOUR pricing policy, and run python3 -m evals.run --run-id b000-<slug>-rules --floor rules with no key. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need, 63 times a quarter. That is the case against the best-fitting scenario (“Your sellers code Approval claimed and Special terms declared honestly, and you only want the seven checks a lookup settles.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A quote whose panels are not the nine this parser knows. src/rules.py matches column layouts by regular expression; a quote exported to a different template parses to zero lines and src/packet.py::check refuses to start rather than scoring an empty document. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; nothing here measures variance between two runs of the same prompt, so read any margin under three checks as unresolved. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-quote-policy. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: it prints all four free floors over the whole corpus in under a second. python3 -m evals.check_labels re-derives the entire answer key with its own parser and its own arithmetic, also free.

A living map of modern AI — kept current every morning