Home › Use Cases › Reconcile one CRM service against what the network inventory actually records
Use caseUC0355
🧪 Use-case kit · runnable

Reconcile one CRM service against what the network inventory actually records

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

⚠︎ This kit has not been graded

0 completed runs, none scored. No committed run is unscored: all six — three free floors, the stub, the scored run and the adversarial arm — carry their own scores block. What IS answered and NOT graded is four FIELDS: backing_circuit, drift and exposure are recomputed by the station and published in both columns for reading rather than for scoring, and why is free text read only by src/refusal.py.

Everything below is coverage, latency and cost. No figure on any of these pages says whether a brief is any good.

The business caseThe problem this solves

A carrier's CRM says what services a customer has and bills for them. The network inventory says what the estate actually records. The two drift apart every day — a circuit is recovered and the record is never closed, a re-provision leaves the build it replaced patched, a migration moves a service and nobody updates the CRM reference, a site address is written two ways — and the reconciliation that would find it is a periodic sample that reads columns. On this corpus that sample's own answer is wrong on 28 of 62 findings and on 26 of the drift sets. Opening one service's reconciliation file, reading the note under every extract row to decide whether the circuit is still there, walking the order trail for a completed provide and a completed cease, checking the operations notes for a migration and whether it actually happened, and deciding by eye whether two written addresses are the same place.

Audience

A carrier's service-assurance or revenue-assurance desk working a reconciliation exception list, and the inventory-data analyst behind it. Whoever decides whether a service goes onto a cease-and-credit queue, a duplicate-build recovery list, or neither. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation file

The corpus is 62 reconciliation file, 0.10 MB (txt 62). It is generated because it has to be. A real CRM-to-inventory extract is a carrier's own customer and network record, and the exact shapes this kit measures — a decommissioned stub the extract still calls IN-SERVICE, a re-provision that left two builds, a migration recorded in prose and never applied to the CRM, the same site written two ways — are the shapes nobody may publish. Generating it also makes the key DERIVABLE rather than written: each file is built as a structure, the blocks are rendered from it, and SRP-2026 is applied to the same structure by the same station every arm gets.

The corpus

  • The 62 reconciliation filegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every account is an invented trading name and every file prints (invented) beside it; account numbers, service ids, circuit ids, record ids, order references, change numbers and site codes are arithmetic on the file index; site addresses are composed from invented buildings and towns and none is a person's. evals/check_labels.py sweeps all 62 files for five families of identifier on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your reconciliation file. That is the whole change — there is no database to migrate.

One reconciliation file, as the model receives itIRC-0001.txt · 1 of 62
SERVICE RECONCILIATION FILE  IRC-0001
Account: ACC-40011 - Brackenhouse Row Logistics (invented)
Service: SRV-80007  Cycle: 2026-07-01 to 2026-07-31  Procedure: SRP-2026
====================================================================================================

THE SERVICE AS THE CRM HOLDS IT
  service id            SRV-80007
  circuit reference     CIR-10013
  product code          ETH-0100-P
  bandwidth kbps        100000
  status                ACTIVE
  activation date       2026-02-21
  cease date            -
  monthly charge        820.00
  site                  Unit 1, Kestrel Park, Bay 2, Marchbank

NETWORK INVENTORY EXTRACT FOR THIS ACCOUNT AND SITE
  INV-0001  CIR-10013  ETH-0100-P       100000  IN-SERVICE  2026-02-21  -           bearer 1000000 kbps, committed rate 100000 kbps
      SITE: Unit 1, Kestrel Park, Loading Bay 2, Marchbank

ORDER TRAIL
  ORD-1000  2026-01-23  provide   completed   CIR-10013  provide order raised from the signed contract

EXTRACT STATUS
  site code             MARCHBANK-01
  extract state         complete
  completed at          2026-07-31 04:12
  records returned      1

RECONCILIATION AS THE PERIODIC SAMPLE REPORTS IT
  finding               ATTRIBUTE-DRIFT
  matched record        INV-0001
  attributes differing  site
  exposure              0.00

OPERATIONS NOTES
  The account's reporting cut moved to month end on 2026-05-01.

====================================================================================================

The outcomeWhat a good result looks like

One service in, one row out: which extract rows have no circuit behind them, whether the migration in the notes actually happened, whether the two addresses are one place, which record backs the service, which attributes disagree, one finding and the monthly figure at risk — with the periodic sample's own answer printed beside it so a reader can see which one moved.

And when it cannot

And what it does when it cannot. On the scored run 61 of 62 replies parsed and NOTHING stopped at the ceiling — the largest reply was 1892 of 4000 output tokens. The one that did not parse (IRC-0026) carried an unescaped newline inside its why and then a second complete JSON object after the first; it is counted WRONG, it stays inside every published denominator, and it was not re-fired. Its second object was in fact the right answer, and taking it would have been the scorer repairing the reply it was about to grade. ⚠︎ ONE REPLY DID REACH THE CEILING, on the adversarial arm's first pass: IRC-0051 stopped at exactly 4000 output tokens with finish_reason length and returned 17 KB of prose no parser could read — 16,891 of its 17,074 characters inside why, the one field in the contract with no bound on it. It was billed in full for $0.002915 and returned nothing, so it was bought AGAIN, alone: its line was dropped from the reply cache and --resume replayed the other fourteen for $0.00. The prompt was NOT touched — a byte changed there would make every already-published call describe a prompt this kit no longer holds. The replacement answered in 268 tokens, refused the injected clause and applied R-4, so the arm now records 0 at the ceiling and 15 of 15 answered.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your inventory's dead records announce themselves — every recovered circuit says "recovered", every duplicate says "duplicate" — the free rules floor, and do not buy a call at all
    A keyword list over the notes reaches everything a keyword can reach, for $0.00, and it already beats the paid arm on cites 55 to 49.
  • Your CRM and your inventory write the same site address in different conventions — abbreviations, a postcode on one side, a descriptive word in front of a bay number — the paid call, and read the address family's own rate first
    site_same is 59 of 62 for the paid arm, 47 for answering "the same place" every time, and 38 for a token-sorted normaliser. A false site drift is a correct service on a work queue, and it is the one field where the money clearly buys something.
  • Your operations notes routinely describe migrations that were planned, quoted or stood down alongside ones that happened — the paid call
    migration is 62 of 62 for the paid arm and 58 for every free arm, and declined_migration_applied is 0 of 62 — the free regex applies all four that never happened because a regex over prose cannot tell them apart.
  • Your extract sometimes does not complete for a site — either — but keep RC-1
    EXTRACT-INCOMPLETE fires on a stated block and every arm gets all 3 of them. It is in the corpus because a reconciliation that cannot say it turns a collector timeout into a customer credit, not because it is hard.

And where nothing here is good enough:

  • You need the citation list to be usable as a work queue — neither, yet
    The paid arm over-cites on 11 services and under-cites on 2. Every over-cited row carries no note at all: it is citing the row its finding is about. A citation list with correct rows in it is a queue nobody can shorten.

At a glanceHow the whole thing runs

76%all five graded fields right, rechecked column
2,010 msp50, end to end
$0.00per 1,000 reconciliation file · google/gemini-3-flash

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own reconciliation files in the same shape and data/services.json with your own CRM rows, keep data/gold.jsonl's five graded fields, and every arm, every grader and every free floor runs unchanged. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per service for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your inventory's dead records announce themselves — every recovered circuit says "recovered", every duplicate says "duplicate"”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An extract that is not fixed-width columns. src/extractsheet.py's row regex is the shape these files print; a CSV, a vendor API export or a graph database needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-inventory-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every screenshot in docs/shots. python3 tools/build_corpus.py --check and python3 -m evals.check_labels both run on a machine with nothing installed — the kit is standard library only and requirements.txt is deliberately empty of packages.

A living map of modern AI — kept current every morning