Home › Use Cases › Decide which of a period's outage minutes the SLA exhibit's own exclusions take out
Use caseUC0514
🧪 Use-case kit · runnable

Decide which of a period's outage minutes the SLA exhibit's own exclusions take out

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An enterprise customer's circuit had a bad month and the carrier's service-assurance desk holds one service-month pack for it: the claim period, the customer's own signed SLA exhibit in both the revisions in force during it, the demarcation record, the planned-maintenance notices, every incident record with its worklog, and whatever the customer has asserted. Somebody has to say, incident by incident, which minutes stay in the chargeable downtime count and which numbered exclusion removes the rest. Two things make that hard and only one of them is arithmetic. The ticket's opened and closed stamps are when a TICKET was open, not when the service was down — across this corpus the stamps alone total 27,742 minutes against the exhibit's 11,503, and 53 of 64 packs land in a different band on the stamps alone. And the exhibit has two revisions: Revision A restores service on ANY path the carrier provides and carries no access-provider exclusion at all, Revision B restores on the working path only and adds E-5. Both print on every pack and the governing one is the one in force at the incident. Reading one service-month pack end to end against the customer's own SLA exhibit: establishing from each worklog when impact actually began and ended rather than when the ticket was open, deciding which numbered exclusion the evidence establishes under the revision in force, then subtracting notified maintenance, taking the union of overlapping incidents, totalling the period, dividing by the period's own minutes and reading the band. It does not replace the claim desk: nothing here calculates, requests, approves, refuses or communicates a credit, names any amount of money, amends the exhibit or touches a ticket.

Audience

The service-assurance desk that has to state a period's chargeable downtime, and the account team that has to be able to show the customer which clause of their own exhibit took each excluded minute out. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual service-month packs

The corpus is 64 service-month packs, 0.29 MB (txt 64). Because the measurement needs a document where the answer is NOT in any column, and a key that knows why. The two facts this kit pays a model for — when impact actually began and ended, and which numbered exclusion the evidence establishes — are in worklog prose and in the revision table, and they are deliberately in conflict with the ticket stamps on 180 of 226 incidents. A real carrier's assurance queue cannot supply that: the packs are somebody's customers, the worklogs are somebody's engineers, and nobody holds a per-incident label saying which clause of which revision took each minute out. Generating it is what makes both halves gradable — and it is also the corpus's own ceiling, which is published rather than hidden: the worklog sentences collapse to 63 distinct skeletons once times, dates and ids are masked.

The corpus

  • The 64 service-month packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your service-month packs. That is the whole change — there is no database to migrate.

One service-month pack, as the model receives itSM-0001.txt · 1 of 64
ENTERPRISE SERVICE ASSURANCE PACK - ONE SERVICE, ONE CLAIM PERIOD, AGAINST THE CUSTOMER'S OWN SLA EXHIBIT

PACK HEADER
  Pack                 SM-0001
  Customer             Aldenmere Logistics Group
  Carrier              Quillhaven Telecom
  Service              CIR-10000
  Product              Dedicated internet access, 500 Mbps
  Master agreement     MSA-4400
  Claim period         2025-06-01 to 2025-06-30 (43200 minutes)
  Compiled by          M. Hallberg, service assurance desk

SLA EXHIBIT (the customer's own signed exhibit, both revisions as the contract file holds them)
   Rev   Effective    Superseded   Notice period   Exclusions carried          Service restored on
   A     2025-01-01   2026-03-31   5 days          E-1, E-2, E-3, E-4          any path, working or protect
   B     2026-04-01   --           10 days         E-1, E-2, E-3, E-4, E-5     the working path only
  Commitment           99.9500 pct availability over the claim period
  E-1                  planned maintenance - minutes falling inside a maintenance window the carrier notified in advance, and only where the notice met the exhibit's own notice period
  E-2                  customer equipment beyond the demarcation point recorded for this service
  E-3                  power or environment at the customer's premises, including a customer standby plant that did not carry the load
  E-4                  a period during which the customer instructed the carrier in writing not to proceed
  E-5                  a fault proved in a third-party access provider's network, in the revisions of the exhibit that carry this exclusion
  Credit table         B1 at or above 99.9000 pct, B2 at or above 99.5000 pct, B3 at or above 99.0000 pct, B4 below that

Abridged — the file continues.

The outcomeWhat a good result looks like

Every incident on the pack carries an exclusion code, an impact interval and — recomputed in pure code from those — a state, its own chargeable minutes and the line of the pack that evidences it; the pack carries the period's chargeable minutes, the availability and the band read off the exhibit's own table. On the published run the shipped column reaches the right exclusion on 223 of 226 incidents, the right interval on 221, the exact chargeable-minute total on 58 of 64 packs and the right band on 62.

And when it cannot

⛔ THE RAW COLUMN'S ARITHMETIC IS NO BETTER THAN FREE CODE AND THAT IS PUBLISHED BESIDE THE HEADLINE. As the reply came back, the model's own chargeable-minute totals are exact on 9 of 64 packs against the free floor of record's 8 — exact two-sided McNemar p = 1.000000, NOT SIGNIFICANT — and its band is BEHIND that floor, 29 of 64 against 36. Everything the shipped column adds above that comes from src/recheck.py, which costs nothing and calls nothing. The money buys the READING; free code does the ARITHMETIC, and a kit that published only the shipped column would be selling arithmetic the model did not do.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your ticketing system already records a service-impact interval separate from the ticket stamps — the free arithmetic alone — src/exhibit.py, no model
    The union, the maintenance subtraction, the period divide and the band lookup are arithmetic on three numbers and are exact for $0.00. On the 18 incidents this corpus settles by column, every arm scores 18 of 18.
  • The interval has to be read out of worklog prose — the call, and then the station
    This is the whole margin: 44 of 49 on the interval families against the free floor's 8, and 200 of 208 on the incidents no column settles.
  • You want the period's chargeable minutes, not the per-incident labels — the call for the readings, always followed by src/recheck.py
    58 of 64 exact with the station against 9 without it, on the same calls. The station is free and calls nothing.
  • You only need the availability band — reconsider the question
    Answering B2 to every pack scores 31 of 64 with no reading at all, against the floor's 36. Five packs separate the two.
  • You need to show the customer the line their own exhibit turns on — the call plus src/citation.py
    220 of 226 admitted, 76 quotes returned, 0 unlocatable — the arm can point at the sentence, not just assert the code.

And where nothing here is good enough:

  • You want a number for a claim or a credit — nothing here
    There is no money anywhere in this kit by construction — 31 forbidden field names asserted against the answer contract at import, and every pack's shipping text grepped for a currency symbol.

At a glanceHow the whole thing runs

99%exclusion pct
2,160 msp50, end to end
$1.78per 1,000 service-month packs · OpenAI GPT-5.6 Luna

Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own service-month packs into data/corpus/ in the same panel layout and the whole free half runs on them: the column reader, the SLA-EX-2026 engine, all four floors, the citation locator, the cap reader and the board. The key is the boundary. Corpus lens →
When is this the wrong choice?Avoid: Do not pay a model to subtract two timestamps. That is the case against the best-fitting scenario (“Your ticketing system already records a service-impact interval separate from the ticket stamps”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack whose panels do not match the fixed layout. src/reader.py is anchored on the panel headings and an incident record that does not match is silently not an incident — the parser returns a pack with fewer incidents rather than raising. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the 9-of-64 raw minute total is really the same as the free floor's 8. The exact two-sided McNemar is p = 1.000000 on 11 discordant packs, so this corpus cannot separate them — that is a published result, not a gap in the harness. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN), one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-21 — r001-sla-downtime. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board and scores every free floor offline: the 64 packs, the derived key, all four floors' committed results, the paid run's committed result and the probe's all ship, and the board replays them at $0.00. Nothing is installed — the kit is standard library end to end. The one control that could spend is disabled with the reason printed beside it, and docs/shots/sla-downtime-empty.png is that board.

A living map of modern AI — kept current every morning