Home › Use Cases › Code a repair order from the technician's story
Use caseUC0188
🧪 Use-case kit · runnable

Code a repair order from the technician's story

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A technician writes the three C's -- concern, cause, correction -- in a free-text box, claims a labour operation code against each line, and is PAID BY THE HOURS THAT CODE CARRIES. Before the claim goes to the manufacturer, a warranty administrator has to decide line by line whether the operation the story supports is the operation that was claimed. The dangerous line is never a typo and it is never arithmetic that does not add up: on this corpus every line FOOTS -- the hours claimed equal the printed allowance of the code claimed, to the tenth, on every line including every wrong one. What is wrong is a PLAUSIBLE code the story does not support: a scope note that excludes the component the cause clause actually names; a member operation claimed beside the superset that already includes it; a causal part that is the part REPLACED rather than the part that FAILED, because the story says the first came out for access; a diagnosis claimed on the same line group as the repair it diagnosed. And one trap runs the other way: every printed field agrees, and one sentence in the shop notes records the part as bench tested, within specification and back in stock, so nothing failed at all. The pass a warranty administrator makes over a repair order before submission: reading each three-C story against the labour-time guide, resolving the technician's shorthand and trade words onto catalogue operations, applying the catalogue's own inclusion and exclusivity rules, and then applying the store's submission standard to decide which lines go back for a better story. It does not replace the signature.

Audience

The dealership warranty administrator who codes and screens repair orders before submission, the service manager who answers for what was submitted, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual repair order (one SYNTHETIC dealer RO: header, customer concern, the technician's three-C story per job, the labour lines and codes claimed, the parts used, the operation catalogue in force for that vehicle, the printed submission standard, the shop notes, and one withheld internal block)

The corpus is 36 repair order (one SYNTHETIC dealer RO: header, customer concern, the technician's three-C story per job, the labour lines and codes claimed, the parts used, the operation catalogue in force for that vehicle, the printed submission standard, the shop notes, and one withheld internal block), 0.31 MB (txt 36). Because the difficulty is REAL and it is not arithmetic. Every line foots, so nothing can be solved by subtraction; the failure under test is always a plausible code the story does not support. And because the trade-language split lets the kit answer the only question a reader actually has -- what does the model buy that a lookup table already had -- as a number on its own denominator rather than as an argument.

The corpus

  • The 36 repair order (one SYNTHETIC dealer RO: header, customer concern, the technician's three-C story per job, the labour lines and codes claimed, the parts used, the operation catalogue in force for that vehicle, the printed submission standard, the shop notes, and one withheld internal block)generated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your repair order (one SYNTHETIC dealer RO: header, customer concern, the technician's three-C story per job, the labour lines and codes claimed, the parts used, the operation catalogue in force for that vehicle, the printed submission standard, the shop notes, and one withheld internal block). That is the whole change — there is no database to migrate.

One repair order (one SYNTHETIC dealer RO: header, customer concern, the technician's three-C story per job, the labour lines and codes claimed, the parts used, the operation catalogue in force for that vehicle, the printed submission standard, the shop notes, and one withheld internal block), as the model receives itRO-0001.txt · 1 of 36
Repair Order
----------------------------------------------------------------
RO Number: RO-0001
Vehicle: 2023 Meridian Cascade 2.0T
VIN: SYNC00007717Z
Odometer: 13,658 mi
Repair month: January
Technician: T-261
Advisor: A-14
Coverage: powertrain and basic warranty ACTIVE for every line on this order

Customer Concern
----------------------------------------------------------------
"Grinding when braking and the pedal pulses at highway speed."
"Clicking on tight turns and a hum that rises with road speed."
"Fan only works on the highest setting and squeals when it does."

Technician Story
----------------------------------------------------------------
[Front brakes]
C:   Grinding when braking and the pedal pulses at highway speed.
CA:  frt brake pads worn to the backing plate, rotors measured within the discard limit.
CO:  R&R front brake pads as an axle set w/ new hardware, bedded in on a rd test.

[Front driveline]
C:   Clicking on tight turns and a hum that rises with road speed.
CA:  LH outer drive axle joint clicking under load w/ play in the joint, boot split w/ it.
CO:  replaced the LH half shaft, cked alignment, rd tested; the CV boot was removed for access and damaged on removal, so it was replaced as well -- it was not defective.

[Climate control]
C:   Fan only works on the highest setting and squeals when it does.
CA:  blower resistor open on the three low speeds, fan motor draw within spec.
CO:  replaced the blower resistor, verified all four blower speeds.

Labour Claimed
----------------------------------------------------------------
Line  Operation as typed by the technician            Hours  Code claimed
1     BRAKE ROTORS FRONT                                1.2  OP-C4104
2     DRIVE AXLE SHAFT                                  1.4  OP-C8130

Abridged — the file continues.

The outcomeWhat a good result looks like

One entry per claimed labour line: the disposition, the operation code this vehicle's own catalogue supports, the part that failed, the allowance that catalogue prints -- and, kept on its own denominator, whether the line goes back to the technician before submission and which of the printed standard's five limbs sends it. Plus, beside every line, what the strongest free floor wrote for nothing.

And when it cannot

A code and an allowance put on a line whose story names no component at all. Measured at 7 of 7 -- the fast tier coded EVERY thin story in this corpus, taking the operation off the parts table instead of the story, while correctly returning all 7 of them to the technician under SS-1 on the question beside it.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your technicians write in catalogue language and your op-code catalogue is small enough to print in one prompt — the free floor (evals/baseline.py, story-lexicon-notes)
    It codes 92.0 pct of catalogue-language cells against the fast tier's 70.7, for nothing, offline, in under a second.
  • Your technicians write in shop shorthand nobody has written down — the model
    On the 17 cells whose trade terms the floor's own table does not carry, the floor scores 23.5 pct and the fast tier 94.1. That gap is the only place on this corpus where a call pays for itself.
  • You want to know which lines to send back to a technician — the model
    97.1 pct return agreement against the best floor's 89.5, all 29 returns caught, the right limb named on every one, and 3 false returns against 10.
  • You need a line the story cannot support to be REFUSED rather than coded — the free floor, and nothing on this page yet
    The fast tier coded 7 of 7 thin stories anyway. The floor refused all 7. This is the axis on which the model is not merely behind but at zero.

At a glanceHow the whole thing runs

76%coding accuracy pct
107,456 msp50, end to end
$45.18per 1,000 repair orders · Google Gemini 3 Flash

Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own repair orders into data/corpus/ in the layout data/SOURCES.md describes -- underlined section headings, a fixed-width labour table ending in the code claimed, a parts table, and the operation catalogue printed with its Includes / Does not include lines. ⚠︎ WHAT DOES NOT TRANSFER. Corpus lens →
When is this the wrong choice?Avoid: The first store that types 'coolant pump' instead of 'water pump' is silently wrong and stays wrong -- the floor scores 23.5 pct on this corpus's 17 such cells -- and nothing in its output says it did not understand the word. That is the case against the best-fitting scenario (“Your technicians write in catalogue language and your op-code catalogue is small enough to print in one prompt”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A real labour-time guide. This catalogue is 24 operations; a manufacturer's is tens of thousands and does not fit in one prompt. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second paid tier. One tier was run through the model seam; that is a spend decision, not a comparison, and nothing here says a deliberating tier would do better or worse. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-27 — r001-ro-opcode. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — git clone, then python3 -m src.app -- no install, no key, no network. The repair orders, the answer key and every committed run file are in the repo, all four free floors are pure Python, and the UI replays what the scored run actually answered straight off the committed result file.

A living map of modern AI — kept current every morning