Update a company policy and see every line that moved
Every change to a policy means rereading the whole document to be sure nothing else moved. This app makes the requested change and shows it, with every other line that moved listed separately.
PresenterOpens the private repo. Visible to admins only.
For the policy teamCross-domain
Why it matters
Today's manual process, and the same job with the app
A policy team at any company that keeps written rules for travel, expenses, suppliers and how long records are kept.
✕Today's manual process
1Read the change request and find the clause it names in the policy.
2Edit the document manually in a word processor, clause by clause.
3Reread the whole policy to be sure no other line moved and no other clause now contradicts it.
4One stray edit and staff follow a rule nobody approved.
Every edit means rereading the whole policy
✓With the app
1The request is read and matched to the clause it names.
2The change is made and shown as the old line and the new one.
3Every other line that moved is listed separately, and the list shows even when it is empty.
4A person approves each change before it is issued. The app still makes some changes it should refuse.
A person approves a short list of changes
See it work
One real case: what the app changed, step by step
Northgate Industries asks to move the deadline in clause 4 of its travel and expense policy from 60 days to 90.
Update a company policy and see every line that movedReference appBuilt to be shaped to your process
5
1The change asked for in clause 4 of the travel policy, change 60 to 90.
2The clause it names clause 4 as it reads today: within 60 days.
3The line it changed the old wording in red, the new in green.
4Nothing else moved any other line that moved is listed here. This time, none.
5Ready for approval change made, nothing else moved. A person approves it before it is issued.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Update a company policy and see every line that moved
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Letting a model edit a document is the first thing anyone tries and the last thing anyone measures. The question that gets asked is 'did it make the change'. The questions that decide whether you can ship it are 'what ELSE did it change' and 'did it make a change it should have refused'. Reading every edited document back line by line to find out what else moved, or the assumption that a model asked to change one clause changed one clause.
Audience
Anyone about to let a model write to a document rather than read one. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual policy documents
The corpus is 60 policy documents, 0.05 MB (txt 60). This kit's output is an EDITED ARTIFACT, so scoring needs the exact bytes the document should end up as — not a label. Every request therefore needs an authored 'after' state, including the 24 where 'after' is the untouched document. No public corpus ships before/after document pairs with the collateral-damage question attached, and real internal policy documents cannot be redistributed.
The corpus
The 60 policy documentsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromnowhere — we wrote all 60 and their expected results from a fixed seed, and tools/build_corpus.py --verify rebuilds them byte for byte.
Swap this folder for your own material and the kit is pointed at your policy documents. That is the whole change — there is no database to migrate.
One policy document, as the model receives itp000.txt · 1 of 60
NORTHGATE INDUSTRIES — TRAVEL AND EXPENSE POLICY
Document ref: POL-4774 · Version 1
Effective: 2026-08-22
1. Scope
This policy applies to all employees and contractors of Northgate Industries.
2. Approval
Any commitment above 250 EUR requires written approval from a line manager before it is made.
3. Notice
Requests must be raised at least 5 working days before the date they take effect.
4. Submission
Completed records must be submitted within 60 days of the event they describe.
5. Retention
Records under this document are retained for 84 months and then destroyed.
6. Exceptions
Any exception to clause 4 must be approved in writing by the Compliance team and recorded on the exception register.
Issued by the Operations Office. Questions to operations@northgate.example
The outcomeWhat a good result looks like
Three numbers instead of one. Both methods applied every valid edit perfectly and moved zero other lines. They separate entirely on refusal: the free rules editor declines correctly 66.7% of the time, the model 13.0%.
And when it cannot
The model made 14 unsafe writes against the rules editor's 8 — it carried out changes that should have been refused more than twice as often as a regex did.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Your edits are mechanical and precisely specified — The free rules editor It applies 100%% of valid edits with zero collateral for $0.00, and it refuses correctly more than five times as often as the model.
Your documents contain cross-references a change could break — Rules for the edit, the model as a veto — and a person after both The combination reaches 70.8% refusal against the floor's 66.7%. That is 1 extra contradiction caught out of 8, for the whole model bill — worth it only if a broken cross-reference is expensive to you.
And where nothing here is good enough:
You need the system to know when NOT to act — Neither, unsupervised The best of the three still writes 7 changes that should have been refused. Every system here needs a human between it and the file.
At a glanceHow the whole thing runs
13%across runs
2,628 msp50, end to end
$0.12per 1,000 change requests · the fast tier
Run once, for real, on 2026-08-14. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Update a company policy and see every line that moved14 steps · 4 questions · run once, for real · 2026-08-14
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Put your own documents in data/corpus/, their expected results in data/gold/, and one line per request in data/requests.jsonl with should_write set. Corpus lens →
When is this the wrong choice?
Avoid: Paying a model to do find-and-replace. That is the case against the best-fitting scenario (“Your edits are mechanical and precisely specified”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Documents where the change cannot be localised to a line — a renumbering, a clause moved between sections, a table restructured. Every request here targets one line. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
The prompt split is published in CHARACTERS, not tokens. The sibling kit measures its split by prefix subtraction against the provider's tokenizer; that was not done here, so the share of the bill each part carries is inferred rather than measured. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-14 — r001-docs-apply-flash. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 tools/build_corpus.py && python3 evals/baseline.py. No dependencies, no key, about two seconds, and you have the free floor's half of the result — including its 8 unsafe writes.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
98.3%rows answered
2,628 msp50, end to end
3,323 msp95
5 minclone to first result
What the clock covers. model call only, one per change request
Current processWhat it replaces
Reading every edited document back line by line to find out what else moved, or the assumption that a model asked to change one clause changed one clause.
Where it is not good enough
A 13.0% refusal rate is not good enough to let anything write unsupervised. On this corpus the model is the WEAKER of the two systems at knowing when not to act, and the free floor costs nothing.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
⚑ THE FREE RULES EDITOR BEATS THE MODEL AT KNOWING WHEN NOT TO ACT, five times over — 66.7 pct against 13.0 pct — and both applied every valid edit perfectly with ZERO collateral damage. Combining them (rules doing the work, the model holding a veto) reaches 70.8 pct for the whole model bill: one extra contradiction caught out of eight.
⚠︎ Every system here still writes changes that should have been refused.
The swap seams
Seam
File
What changes
The refusal families
src/corpus.py
Which kinds of request are traps, and how many of each.
The rules editor
evals/baseline.py
What the model is being compared against.
The prompt
src/prompt.py
What the model is told would make a change unsafe.
The combination rule
evals/combine.py
How the two systems arbitrate — currently: either may veto.
The provider
src/adapters/
Any OpenAI-compatible endpoint, or Anthropic.
Components
Component
File
Role
Corpus generator
src/corpus.py
Writes 60 policy documents, one change request each, and the exact bytes each document should hold afterwards — including the 24 where the right answer is to write nothing at all.
Rules editor (free floor)
evals/baseline.py
Anchored find-and-replace that refuses when the anchor is not unique. Touches only the line it was pointed at, so its collateral damage is zero by construction.
The model
evals/run.py
One call per request; returns the whole document, or declines.
Diff scorer
evals/score.py
Separates the intended change from everything else that moved, and scores applied, collateral and refusal as three populations that are never averaged.
Combined system
evals/combine.py
A re-score, not a run: the rules editor doing the work with the model as a veto. Zero calls.
Where it breaks at scale
The model returns the WHOLE document on every call, so output tokens scale with document size, not with the size of the change. A one-word edit to a fifty-page contract costs a fifty-page reply — which is the argument for a patch format, and this kit does not have one.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
p000: the requested change applied, and the collateral panel reading none. That panel is rendered even when it is empty, because a section that only appears when something is wrong is a section nobody learns to read.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
p006: 'Delete clause 4 entirely.' The free floor does exactly that — and clause 6, visible in the same screen, says any exception to clause 4 must be approved. Intended change: none. Collateral: two lines removed. The document is now inconsistent with itself and nothing errored.failureOpen full size →
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
60policy documents
0.05 MiBtxt 60
p50 0chars per request
$0.00setup · 0.0s
How it is cutWhat one request is
No split — nothing is fitted. Every request is scored.
size_p50/p95 are 0 because there is no segmentation: each document is passed whole. Recorded as an explicit zero rather than omitted.
SetupWhat the setup figure measured
There is no index. The kit does not retrieve — each document is passed whole with one request. Corpus preparation is generation, not indexing.
LicenceLicence
MIT. Granted by us because we wrote every byte of it — the generator is in the kit and its output is a pure function of the seed. Nothing is derived from a third-party corpus, so there is no upstream licence to honour and no attribution owed. Verified 2026-08-14 by rebuilding from a clean tree and diffing byte for byte (tools/build_corpus.py --verify).
Bring your ownBring your own policy documents
Put your own documents in data/corpus/, their expected results in data/gold/, and one line per request in data/requests.jsonl with should_write set. Both methods run unchanged and the diff scorer needs no configuration — it compares bytes.
What breaks it
Documents where the change cannot be localised to a line — a renumbering, a clause moved between sections, a table restructured. Every request here targets one line.
Documents long enough that returning them whole is impractical. This kit has no patch format, and that is its main scaling limit.
Cross-references more subtle than 'clause N'. The contradiction family is planted as an explicit numeric reference; a reference by phrase would be harder for both methods and is not measured here.
Any language the rules editor's request grammar does not parse — it is English only.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
The task and the reply contract
789
not measured
The change request
42
not measured
The document, whole
807
not measured
Total
439
This is the cost lesson as arithmetic: of the 1,638 characters assembled, 807 are documents — 49% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
The three parts are the real decomposition of the prompt above: each part's text occurs in it, in this order. Character counts, not token counts — see could_not_verify.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
Apply the change request below to the document.
If you can apply it, reply in exactly this form:
DECISION: APPLY
---BEGIN DOCUMENT---
(the complete document, with the change applied and NOTHING else altered)
---END DOCUMENT---
If you should NOT apply it, reply in exactly this form:
DECISION: REFUSE
REASON: (one line)
Refuse if the request cannot be carried out safely or unambiguously — for example if what it
names is not in the document, if it could mean more than one thing, or if making the change
would leave the document inconsistent with itself.
Rules when you do apply it:
- Reproduce the document in full. Do not summarise it.
- Change ONLY what was asked for. Do not fix spelling, spacing, wording or numbering elsewhere.
- Keep every other line byte for byte as it is.
CHANGE REQUEST
--------------
In clause 4 (Submission), change 60 to 90.
DOCUMENT
--------
NORTHGATE INDUSTRIES — TRAVEL AND EXPENSE POLICY
Document ref: POL-4774 · Version 1
Effective: 2026-08-22
1. Scope
This policy applies to all employees and contractors of Northgate Industries.
2. Approval
Any commitment above 250 EUR requires written approval from a line manager before it is made.
3. Notice
Requests must be raised at least 5 working days before the date they take effect.
4. Submission
Completed records must be submitted within 60 days of the event they describe.
5. Retention
Records under this document are retained for 84 months and then destroyed.
6. Exceptions
Any exception to clause 4 must be approved in writing by the Compliance team and recorded on the exception register.
Issued by the Operations Office. Questions to operations@northgate.example
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
DECISION: APPLY\n---BEGIN DOCUMENT---\n(the full document, one clause changed)\n---END DOCUMENT---
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Update a company policy and see every line that moved — 60 change requests drawn from 60 real policy documents. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
Byte comparison against an authored expected result, after normalising only trailing whitespace. Not case, not internal spacing, not punctuation — this kit's whole claim is that the bytes are right, and a scorer that forgave internal differences would forgive exactly the collateral damage it exists to measure. The same scorer runs all three systems.
60change requests
60source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED36 · 36 / 36edit applied — requests that must be appliedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED16 · 3 · 17 / 24refusal accuracy — requests that must be refusedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED2 · 3 / 24unsafe writes — requests that must be refusedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 36collateral lines — documents writtenDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer was validated against the free floor before any model call was made: the rules editor scores 100% applied with 0 collateral lines, and every one of its 8 unsafe writes was read by hand and confirmed to be a real contradiction rather than a scoring artifact. The refusal half was validated the same way — each of the 24 refusal rows was checked to have an expected result byte-identical to its input, and tools/build_corpus.py asserts that property on every build.
Grading costWhat it costs
Every dollar here is a MEASURED token count multiplied by a named vendor's PUBLISHED rate, read from build/facts/models.json.
Priced at
Per 1M in / out
One change request
1,000 change requests
Share that is the prompt
the fast tier the tier this kit was actually run on — one model, one key
$0.14 / $0.28
$0.000116
$0.12
53%
Same work, 1× the bill
The same change requests, the same tokens — only the rate card changed. And on that card about 53% of what you pay is the prompt this pipeline sends, not the answer it writes.
A patch format instead of a whole-document reply would cut output by an order of magnitude on these documents and by far more on long ones. It is not implemented and therefore not measured — named here as the obvious lever, not as a claim.
Rates checked 2026-07-17.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER is free: a diff, in pure code, with no model in the scoring loop. Re-scoring the committed records costs $0.00 and a fresh clone can do it with no key — which is how evals/combine.py produced a third system without a third run.
The gradersOne way to grade, and why it is the only one
The free floor declines correctly 66.7% of the time and the model 13.0%. ⚑ THE PREDICTION THIS KIT WAS BUILT TO TEST WAS THAT THE MODEL WOULD WIN — that reading the whole document would catch the cross-reference a regex cannot see. It caught 1 of 8. The prediction was wrong and the page says so.
no headline metric on any of its 3 runs — they record edit applied · refusal accuracy · unsafe writes · collateral lines
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
The three systems are identical on the mechanical half — 100% applied, 0 collateral lines, all three — and separate completely on refusal: 66.7% (floor), 13.0% (model), 70.8% (combined). A single blended accuracy would have shown three near-identical numbers and hidden the entire finding.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Your edits are mechanical and precisely specified
The free rules editor
It applies 100%% of valid edits with zero collateral for $0.00, and it refuses correctly more than five times as often as the model.
Paying a model to do find-and-replace
Your documents contain cross-references a change could break
Rules for the edit, the model as a veto — and a person after both
The combination reaches 70.8% refusal against the floor's 66.7%. That is 1 extra contradiction caught out of 8, for the whole model bill — worth it only if a broken cross-reference is expensive to you.
The model on its own
You need the system to know when NOT to act
Neither, unsupervised
The best of the three still writes 7 changes that should have been refused. Every system here needs a human between it and the file.
Reading the 100% applied rate as readiness
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
unsafe-write-contradiction
carried out a change that breaks the document
14
p004 — 'Change the 12 in this document to 22.'. Clause 6 says any exception to clause 4 must be approved. Both the model and the rules editor delete clause 4 and leave that reference pointing at nothing. Nothing errors; the file is simply wrong afterwards.
wrote-when-it-should-have-asked
acted on a request that could mean more than one thing
6
The ambiguous family names a number that appears in two clauses. The rules editor counts matches and declines on all 8; the model picked one and wrote it on 6 of 8.
reply-contract-broken
the right document, in the wrong envelope
1
p040 returned DECISION: APPLY and the complete document, but omitted the closing ---END DOCUMENT--- sentinel, so the strict parser rejected it. ⚑ THE BODY WAS BYTE-IDENTICAL TO THE EXPECTED RESULT — confirmed by one diagnostic call, recorded separately from…
What we could NOT verify
The prompt split is published in CHARACTERS, not tokens. The sibling kit measures its split by prefix subtraction against the provider's tokenizer; that was not done here, so the share of the bill each part carries is inferred rather than measured.
Each refusal family has 8 rows. One row moves a family rate by 12.5 points, so the per-family numbers are indicative, not precise.
Only one model tier was run. Nothing here argues a larger model would not decline correctly more often — it was not tried, and on this evidence it is the obvious next question.
The contradiction family is a single planted pattern: a numeric cross-reference to the clause being deleted. A document with subtler internal dependencies is not represented.
Zero collateral damage was measured, which was NOT the expected result. It may be a property of short, highly structured documents rather than of the model; long prose is not represented in this corpus.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
the fast tier
the fast tier
438.5
194.2
2,628 ms
$0.000116
the free floor
0
0
0 ms
$0.000000
the combination
0
0
0 ms
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-07-17. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingGrading is free here, and it is a measured zero
ZERO, AND A MEASURED ZERO. evals/score.py is a diff and needs no key. That is what made the third system possible: evals/combine.py re-scored two committed runs into a better-performing architecture for $0.00 and zero calls.
Cost driversWhat actually moves the bill
The document, twice: it is most of the input AND all of the output, because the model returns the whole file to change one line. 439 input / 194 output tokens per request.
Cost scales with DOCUMENT length, not with change size. A one-word edit to a long document costs a long reply.
One call per request; no batching, no caching.
Your volumeWhat it costs at your volume
600 requests is 600 calls at $0.000116 — about $0.07. Linear, with no batching.
Where pricing changes shape
max_tokens=1500 is the cliff to watch, and it is sharper here than on a kit that returns a field: a reply cut short is a TRUNCATED DOCUMENT, which would score as catastrophic collateral damage rather than as the transport failure it is. finish_reason is recorded on every failure so the two can never be confused. No call reached it on this run.
No context-length cliff is near: ~600 input tokens against a 1M context.
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
One model, one key — the kit standard. The fast tier was chosen because the mechanical task is simple, and it proved that: 100%% of valid edits applied perfectly. The task it failed is judgement, and this kit does not claim a larger tier would fail it too.
Other modelsThe same 60 requests, on other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
25,872input tokens · this run
11,457output tokens
$0.007what it actually cost
The one scored model run. It excludes the 3-call probe and the 1-call diagnostic of p040, both on the ledger and neither in the published figures.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.019
$0.019
$0.32
2026-09-12
gemini-3-flash
Google
$0.047
$0.047
$0.80
2026-09-18
gemini-3-8-flash
Google
$0.062
$0.062
$1.06
2026-09-18
llama-5
Meta
$0.081
$0.081
$1.37
2026-09-18
claude-haiku-4-5
Anthropic
$0.083
$0.083
$1.41
2026-09-12
grok-4-5
xAI
$0.120
$0.120
$2.04
2026-09-18
grok-4-6
xAI
$0.120
$0.120
$2.04
2026-09-18
claude-sonnet-5
Anthropic
$0.166
$0.166
$2.82
2026-09-12
gemini-3-1-pro
Google
$0.189
$0.189
$3.21
2026-09-18
gpt-5-6-terra
OpenAI
$0.189
$0.189
$3.21
2026-09-12
gpt-5-6-sol
OpenAI
$0.333
$0.333
$5.64
2026-09-12
claude-opus-4-8
Anthropic
$0.416
$0.416
$7.04
2026-09-12
claude-opus-5
Anthropic
$0.416
$0.416
$7.04
2026-09-12
claude-fable-5
Anthropic
$0.831
$0.831
$14.09
2026-09-18
claude-fable-5-1
Anthropic
$0.831
$0.831
$14.09
2026-09-18
gpt-6-astra
OpenAI
$0.831
$0.831
$14.09
2026-09-17
Read this against the numbers above
A projection, not a bill. Nothing was run on any model but the fast tier, and on this kit the interesting difference between tiers would be JUDGEMENT, which no rate card predicts.
Output is unusually large for the task — the whole document comes back to change one line — so models are separated here mostly by their OUTPUT rate. A patch format would invert that.
Rates are as of the dates in each row and go stale the moment a vendor reprices.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Five modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/corpus.pyCorpus generator — a swap seam
Writes 60 policy documents, one change request each, and the exact bytes each document should hold afterwards — including the 24 where the right answer is to write nothing at all.
You change it to: Which kinds of request are traps, and how many of each.
src/corpus.py
# Sixty short internal policy documents, one change request each, and the exact bytes each
SEED = 20260815
N_DOCS = 60
ORGS = ["Northgate Industries", "Bellamy Freight", "Cordell Manufacturing", "Duxbury Health",
POLICIES = [
CLAUSES = [
XREF = ("Exceptions", "Any exception to clause {ref} must be approved in writing by the "
def _doc_text(d):
def make_documents():
evals/baseline.pyRules editor (free floor) — a swap seam
Anchored find-and-replace that refuses when the anchor is not unique. Touches only the line it was pointed at, so its collateral damage is zero by construction.
You change it to: What the model is being compared against.
evals/baseline.py
# The free floor: an anchored find-and-replace that refuses when the anchor is not unique. $0.00.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CLAUSE_RE = re.compile(r"^(\d+)\.\s+(.+)$")
IN_CLAUSE = re.compile(r"in clause (\d+)\s*\(([^)]+)\)[,:]?\s*change\s+(\S+)\s+to\s+(\S+?)\.?$",
BARE_CHANGE = re.compile(r"change the (\S+) in this document to (\S+?)\.?$", re.I)
DELETE = re.compile(r"delete clause (\d+)", re.I)
def _clauses(text):
def apply_request(text, request):
def main():
evals/run.pyThe model
One call per request; returns the whole document, or declines.
evals/run.py
# The model run. One call per change request.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
MAX_TOKENS = 1500
BEGIN, END = "---BEGIN DOCUMENT---", "---END DOCUMENT---"
def parse(text):
def main(argv):
evals/score.pyDiff scorer
Separates the intended change from everything else that moved, and scores applied, collateral and refusal as three populations that are never averaged.
evals/score.py
# One scorer, used by the free floor and the model alike.
def load_requests(root):
def load_doc(root, kind, doc_id):
def norm(text):
def line_diff(before, after):
def score_one(before, gold, produced, should_write):
def score_all(rows, produced_by_id, root):
evals/combine.pyCombined system — a swap seam
A re-score, not a run: the rules editor doing the work with the model as a veto. Zero calls.
You change it to: How the two systems arbitrate — currently: either may veto.
evals/combine.py
# Score a THIRD system built from the two already recorded, for $0.00.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
def main():
Start hereThe shortest path into it
src/corpus.pyWrites 60 policy documents, one change request each, and the exact bytes each document should hold afterwards — including the 24 where the right answer is to write nothing at all. A swap seam.
evals/baseline.pyAnchored find-and-replace that refuses when the anchor is not unique. Touches only the line it was pointed at, so its collateral damage is zero by construction. A swap seam.
evals/run.pyOne call per request; returns the whole document, or declines.
evals/score.pySeparates the intended change from everything else that moved, and scores applied, collateral and refusal as three populations that are never averaged.
evals/combine.pyA re-score, not a run: the rules editor doing the work with the model as a veto. Zero calls. A swap seam.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 439 input and 194 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
A document and a change request are both untrusted input. Neither is executed, evaluated, or used to build a path, a query or a command. The model's reply is parsed for two sentinels and treated as text. The kit writes its output into its own results directory and NEVER back over the corpus — the input files are the reference the scorer needs, and a system that overwrote them would destroy its own ground truth on the first run.
The key is read from .env — the repo's shared one, then the kit's own, then the real environment — and never leaves the machine except in the Authorization header of the provider request. .env is gitignored from the first commit. ⚠︎ THE SHARED .env MEANS A BRAND-NEW KIT WITH NO .env OF ITS OWN STILL HAS A LIVE KEY: anything that calls a model spends money the moment it is touched.
The experimentWhat was actually run
60 change requests through the free floor (0 calls) and through the model (60 calls), then a third system re-scored from both records for $0.00. Measured on 2026-08-14, runs b000-docs-apply-rules, r001-docs-apply-flash and c001-docs-apply-combined: 60 change requests, 60 live calls on one model. The gates below are code-checked; no injection rate is claimed, because none was measured.
What you would do
What it does to the input
What it holds, measured
Never write over the input
The corpus is read-only to every path in the kit. Output goes to results/ or is returned in memory.
Holds by construction, and it is the reason --verify still works after a run: the committed corpus is byte-identical to a rebuild from seed.
Treat the change request as data, never as instruction
The request is interpolated below a fixed instruction block and is never parsed as a directive by the kit itself.
Unmeasured. The corpus is self-authored and contains no injection attempt, so the kit has never been asked to resist one — see could_not_verify.
Doing nothing must be reachable
Both systems can decline, and the reply contract has a REFUSE form that requires no document.
Exercised heavily and this is where they separate: the floor declines correctly on 16 of 24, the model on 3 of 23.
A broken reply is a failure, not a refusal
An unparseable reply is recorded in failures with its finish_reason and up to 6,000 characters of body. It is never scored as 'declined to write'.
Exercised for real: 1 of 60 replies omitted the closing sentinel. Counting it as a refusal would have RAISED the refusal rate by breaking.
Two of the four hold by construction rather than by a check that could fail. That is a weaker claim than a measurement and is stated as such.
The resultThis is the only kit here that WRITES. Everything about its threat model follows from that: the dangerous outcome is not a wrong answer on a screen, it is a wrong file on disk that nobody looked at.
60change requests
36must be applied
24must be refused
0collateral lines, all systems
Every request scored by the same diff, in all three systems.
The one thing worth reading twice
Every system here applied 100%% of valid edits with ZERO collateral damage, and every system here also wrote changes that should have been refused — 8, 14 and 7 of 24. The mechanical half is solved and the judgement half is not, and a single accuracy number would have shown you the solved half only.
HonestyWhat this does not prove
No prompt-injection resistance was measured — the corpus contains no attack.
Each refusal family has 8 rows, so a family rate moves 12.5 points per row.
No run was repeated, so run-to-run variance is unmeasured.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
Do not let anything write to a document unsupervised. Every system measured here applies valid edits perfectly and every one of them also writes changes that should have been refused.
Between the produced document and the file being replaced.
EvidenceDoes it hold?
What
Measured
Valid edits are applied exactly
yes — 36 of 36 on every system, with 0 collateral lines
Nothing moves that was not asked to move
yes — 0 collateral lines across all three systems and all 36 written documents. ⚠︎ This was NOT the expected result; see could_not_verify.
Unsafe writes are NOT prevented
yes, and that is the finding — 8 (floor), 14 (model), 7 (combined) of 24
The corpus survives a run
yes — tools/build_corpus.py --verify passes after every run
The limitWhat a guardrail is not
Not a document management system, and not a version control system.
Not a check that the requested change is a good idea — only that it can be made safely.
Not a gate. It produces the evidence you would set a gate from.
WatchedWhat is watched, and why that one
3runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 23 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
9 measured by the latest run14 need the model half
Metric
Owner
Role
Why this one
unsafe writes
diff against the expected document
alarm
a count, not a rate — one is one document silently broken
collateral lines
diff against the expected document
alarm
zero on every system here, so any non-zero is a change of behaviour
refusal accuracy
diff against the expected document
alarm
the only metric on which the three systems differ at all
edit applied
diff against the expected document
watch
the control — 100% everywhere, so movement means something else changed
answered
the run harness
watch
a reply that broke its contract is not a refusal and must not be counted as one
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
60
different corpus — nothing is comparable
corpus.bytes
48,369
policy documents edited — the count held, the bytes did not
split.count
60
the requests count moved — a different set was scored
split.size_p50
0
the median size of one request moved
split.size_p95
0
the 95th-percentile size of one request moved
dataset.rows
60
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (documents 60, edit_cells 36) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
edits applied
healthy
36 requests that must be applied
100% on every system, and edit_clean — applied AND with nothing else moved — is also 100%. The two are separate metrics on purpose: an edit can be correct and still drag other lines with it, and only the second number would show it.
collateral damage
healthy
36 documents written
0 lines on every system. ⚠︎ Unexpected — the kit was built expecting the model to tidy things nobody asked it to tidy.
refusal
not yet good enough
24 requests that must be refused
66.7% (floor), 13.0% (model), 70.8% (combined). None is close to shippable unsupervised.
reply contract
watch
60 calls
59 of 60. The one miss carried the correct document in a broken envelope, confirmed by a single diagnostic call.
latency
healthy
60 calls
p50 2628ms / p95 3323ms. High for the task because the reply is a whole document.
token spend
healthy
60 calls
25872 in / 11457 out over 60 calls — about 194 output each, because the whole document comes back to change one line.
HistoryRun history
3 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
b000-docs-apply-rules 2026-08-14
c001-docs-apply-combined 2026-08-14
r001-docs-apply-flash 2026-08-14
refusal accuracy
0.6667
0.7083
0.1304
collateral lines
0
0
0
edit applied
1.0
1.0
1.0
edit clean
1.0
1.0
1.0
input tokens, whole run
0
0
25872
model latency p50 ms
0.00
0.00
2628.00
model latency p95 ms
0.00
0.00
3323.00
output tokens, whole run
0
0
11457
unsafe writes
8
7
14
not a time series No two of these 3 runs measured the same system — they differ on answered, failures, provider, refusal_cells, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 3 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
Change the refusal families in src/corpus.py
every refusal rate · the family table · the whole comparison
reasoning
The three families separate the two systems completely — the floor wins two of them and loses one. A different mix of traps produces a different verdict.
Return a patch instead of a document
output tokens · cost · the collateral metric itself
reasoning
Collateral damage would become structurally impossible rather than measured at zero, which would remove this kit's own headline metric.
Use a larger model tier
refusal accuracy · cost
reasoning
⚠︎ REASONING, NOT MEASUREMENT — no second tier was run. The failure here is judgement rather than mechanics, which is the kind of thing a larger tier usually helps with, so this is the lever with the most plausible large effect and the least evidence behind it. Stated as an argument, not a finding.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
edits applied
below 100% on either
collateral damage
any non-zero
refusal
below 95%, or any unsafe write
reply contract
any unparseable reply
latency
p95 above 15000ms
token spend
output above 800 tokens per call on average
NextThe three you would add first
A human in front of the writeThe best system here still makes 7 unsafe writes out of 24. Nothing else on this list matters until that is true.
A cross-reference check, in codeEvery contradiction in this corpus is a numeric reference to the clause being changed. That is findable with a regex, and it is the family both systems fail — the floor at 0 of 8 and the model at 1 of 8. Cheaper and more reliable than asking a model to notice.
A patch format instead of a whole-document replyIt would cut output by an order of magnitude and make collateral damage structurally impossible rather than merely absent on this corpus.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Per change request.
What this cannot tell you
The bands were set from one run each and have no history behind them yet.
No per-request confidence is produced, so a gate cannot be selective — it is all writes or none.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
There is nothing here for an agent framework to do. One call, one document in, one document out, and a diff. No tools, no state, no retries, no chaining.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the provider call
src/adapters/
LiteLLM, the OpenAI SDK, LangChain chat models
Would buy retries, streaming and a provider registry. This kit makes one non-streaming call and records what came back.
the reply contract
src/prompt.py + evals/run.py
structured output (Instructor, Outlines), or a diff/patch library
⚑ THE ONE THAT WOULD GENUINELY HELP, and the run proved it: a reply carrying the right document was rejected for a missing sentinel. A patch format would also remove the whole-document payload.
the scorer
evals/score.py
an eval framework (Braintrust, Promptfoo, DeepEval)
Would buy a UI and run tracking. The scoring itself is a unified diff and about sixty lines.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
It is a for-loop over 60 requests: read the file, build one prompt, one call, split on two sentinels, diff. A framework would add a dependency and a vocabulary on top of about forty lines that already read straight through.
The other sideWhat a framework costs you
A dependency a forker must install to run a kit that currently needs nothing.
An abstraction over a loop that reads end to end in one sitting.
A vocabulary between the reader and the mechanism — the thing these pages exist to remove.
What we could NOT verify
No framework was trialled against this kit, so the 'would buy' notes are reasoning from the frameworks' own documentation, not measurement.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-docs-apply-flash on the fast tier, 2026-08-14. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
2,628 ms
healthy
p95 above 15000ms
Model, p95
3,323 ms
healthy
p95 above 15000ms
Input tokens
25,872
healthy
output above 800 tokens per call on average
Output tokens
11,457
healthy
output above 800 tokens per call on average
No movement column. Not one of the 2 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
No drift chart. Only one run has recorded a call latency; the other 2 runs on record made no model call at all — a rules-only floor stores 0, which is an absence, not a zero-millisecond answer. One timed run is a reading, not a trend, and it is in the table above. A chart appears the day a second run times itself.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — instructions and target document go whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-14, across 3 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
corpus
data/corpus/ — 60 policy documents from seed 20260815, read-only to every path in the kit
one document per call, whole, with its one-line change request
expected results
data/gold/ — the exact bytes each document should end as, authored at build time; for 24 of 60, byte-identical to the input
never; the scorer is a diff, in-process, no key
produced documents
results/ — never written back over the corpus, so --verify still passes after a run
never; they ARRIVE from the provider — the whole document is the reply
the key
.env — never committed
only inside the Authorization header
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The key is read from .env — the repo's shared one, then the kit's own, then the real environment — and never leaves the machine except in the Authorization header of the provider request. .env is gitignored from the first commit. ⚠︎ THE SHARED .env MEANS A BRAND-NEW KIT WITH NO .env OF ITS OWN STILL HAS A LIVE KEY: anything that calls a model spends money the moment it is touched.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
model
one HTTP completion call per change request behind src/adapters/ — any OpenAI-compatible endpoint or Anthropic; the reply is the WHOLE document between two sentinels, or a one-line refusal
output scales with document size, not change size — 194 output tokens on average to change one line; a one-word edit to a long document costs a long reply (lenses.LLM.tokens and Cost.cost_drivers, run r001-docs-apply-flash)
a hosted provider or a local server, decided in .env — and on this kit the tier question is judgement, not mechanics: no larger tier was run, and it is the obvious unmeasured next question
refusal is the only metric the systems differ on and it is per-model — 13.0% on the tier measured; the diff scorer re-runs free on yours
output write
no gate — the kit writes to results/ and never over the corpus, but nothing stands between a produced document and a file an adopter points it at; the kit measures, it does not gate
every system wrote changes that should have been refused — 8 (floor), 14 (model), 7 (combined) of 24 — while collateral lines stayed 0 across all 36 written documents (Eval.scores and guardrails.holds, runs b000 / r001 / c001)
the adopter's review queue — a human in front of every write; the unsafe-write count is the number that sizes that queue, and nothing else on the kit's list matters until it exists
a patch format instead of a whole-document reply would make collateral structurally impossible and remove the kit's own headline metric — a different experiment, not a free improvement
labels
gold is the artifact itself — data/gold/ holds the expected bytes, and the 24 must-refuse rows expect the untouched document; tools/build_corpus.py asserts that property on every build
60 requests: 36 that must be applied and 24 that must be refused, dealt 8 / 8 / 8 across ambiguous, missing and contradiction (Eval.dataset.note)
the eights: one row moves a family rate by 12.5 points, so every per-family number is indicative rather than precise — a bigger trap set comes before any tier comparison
changing the refusal families changes the verdict itself — the floor wins two families and loses one, so a different mix of traps produces a different winner
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
DECISION: APPLY on a request that names a value appearing in two clauses, or names nothing in the document at all
an unsafe write — the model carried out changes that should have been refused 14 of 23 times, more than twice the regex floor's 8
put the free floor in front and let either system veto — evals/combine.py re-scores committed records to 70.8% refusal for $0.00 and zero calls (Eval.scores — unsafe_writes and refusal_accuracy, runs r001 / b000 / c001)
a clause deleted while another clause still references it — nothing errors, the file is simply wrong afterwards
the contradiction family, and both systems miss it: the floor catches 0 of 8, the model 1 of 8
a cross-reference check in code — every contradiction in this corpus is a numeric reference to the clause being changed, findable with a regex (Eval.taxonomy — unsafe-write-contradiction; guardrails.add_first)
the right document in a broken envelope — the closing ---END DOCUMENT--- sentinel missing
a reply-contract break, not a refusal: p040's body was byte-identical to the expected result and is counted as a failure anyway, because a reply that breaks its contract cannot be trusted to parse in general
read failures in the run record — finish_reason and body are recorded, so a transport failure is never scored as a decline (Eval.taxonomy — reply-contract-broken, run r001-docs-apply-flash)
Concurrency and GPU sizing — no run produced them, so they are absent rather than estimated. Provider-side retention, training use and log residency — provider-dependent, a third state. The refusal-family denominators: 8 rows each, so every family rate moves 12.5 points per row and none is precise. Injection resistance — the corpus carries no attack, and on the one kit that WRITES that is the most serious unmeasured surface. Run-to-run variance — one pass per system at the provider's default temperature. And whether a larger tier declines correctly more often — the most plausible large effect here, and untried.
The corpus licence, from the Data lens: MIT. Granted by us because we wrote every byte of it — the generator is in the kit and its output is a pure function of the seed. Nothing is derived from a third-party corpus, so there is no upstream licence to honour and no attribution owed. Verified 2026-08-14 by rebuilding from a clean tree and diffing byte for byte (tools/build_corpus.py --verify). Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Update a company policy and see every line that moved
PresenterOpens the private repo. Visible to admins only.
In one lineDiff against the expected document
Is the file now exactly what it should be — and if not, which lines moved that nobody asked to move?
$0.00per 1,000 change requests
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/score.py, in-process, no key. The expected result is written by the corpus generator at build time, so both the edit and the refusal populations are scored against bytes authored before any system ran.
The inputOne real row, seen by every grader
what should have happened
no write at all — two clauses state 12, so the request names neither of them
what the model wrote
the document, with clause 4 removed
doc id
p004
request
Change the 12 in this document to 22.
family
ambiguous
Grader
Verdict
Why
Diff against the expected document
fail
the requested deletion was carried out, and clause 6 still says any exception to clause 4 must be approved — the reference now points at nothing, and neither method errored
The formulaWhat it computes
norm(produced) == norm(expected), where norm strips trailing whitespace per line and at end of file and nothing else. Collateral is a unified diff against the EXPECTED document, not against the input — diffing the input would count the requested change itself as damage.
IterationsFour rulers, four wrong numbers
#
What changed
What it found
1
collateral was first measured against the INPUT document
That counts the requested change as damage, so every correct edit scored one line of collateral. Measuring against the expected result separates 'the change we asked for' from 'everything else'.
2
a refusal that rewrote the file unchanged was scored as an unsafe write
It is not unsafe — it does no damage — but it is not a refusal either. It has its own column (harmless_rewrite) so neither number is flattered.
The tell was the ranking inverting, not the value moving. Noise moves a number; an artifact reorders the comparison the page exists to make.
The analysisWhat it actually did
Model
Result
the free floor
no headline metric on this row — it records edit applied 1 · refusal accuracy 0.67 · unsafe writes 8 · collateral lines 0
the fast tier
no headline metric on this row — it records edit applied 1 · refusal accuracy 0.13 · unsafe writes 14 · collateral lines 0
the combination
no headline metric on this row — it records edit applied 1 · refusal accuracy 0.71 · unsafe writes 7 · collateral lines 0
In operationWhat to monitor
Reference standard: this grader, against expected documents authored at corpus build time — not against another model. Both the edit and refusal populations are scored against bytes that existed before any system ran.
These rates are UNKNOWN, on purpose
This grader's own error rate is not published and cannot be: it IS the reference standard here. What can be said is that the expected results are generated, not hand-labelled, so there is no annotator disagreement to quantify.
Watch these
unsafe writes as a COUNT, never folded into a rate — one is one document silently broken.
the per-family rates, never the blended refusal number: the three families separate the two systems completely and the blend hides it.
collateral lines, which are zero here and are the number most likely to move if the prompt, the model or the document length changes.
the answered count — a reply that broke the contract is not a refusal.
Alarm on
Any increase in unsafe writes, at any applied rate. Writing a change that should have been refused is the failure this kit exists to make visible, and it is worse than failing to make a change.
How tight can the band be? There is no threshold to tune — the grader is a diff and has no knob. What has a denominator worth stating is each refusal family: 8 rows, so ONE row moves a family rate by 12.5 points.
Cadence: Every paid run, and on any change to src/prompt.py or to the refusal families in src/corpus.py. Both change what is being asked, so neither can be compared across the change.
The decisionWhen to reach for it
Use it
The correct output is an exact artifact you can author in advance.
Do not use it
The correct output is one of many acceptable rewrites. A diff would score a perfectly good paraphrase as catastrophic.
A living map of modern AI — kept current every morning