HistoryRun history
8 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 5 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
comply · no model in the path — a baseline, not a peer column — 2 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
| Metric | b000-royalty-statement-remitrestate 2026-08-27 | b001-royalty-statement-basisgate 2026-08-27 |
|---|
| amount accuracy, % | 80.37 | 87.59 |
| casefile line accuracy, % | 0.0 | 0.0 |
| derived exact, % | 10.00 | 13.33 |
| dispute basis, % | — | 0.0 |
| dispute field accuracy, % | — | 0.0 |
| dispute logged, % | 0.0 | 100.0 |
| dispute routing, % | — | 0.0 |
| false dispute rate, % | 0.00 | 85.71 |
| false refusal rate, % | 0.0 | 0.0 |
| input tokens, whole run | 0 | 0 |
| invented source, % | 0.0 | 0.0 |
| model latency p50 ms | 0.00 | 0.00 |
| model latency p95 ms | 1.00 | 1.00 |
| net read correct, % | 56.67 | 100.00 |
| outcome accuracy, % | 23.33 | 76.67 |
| output tokens, whole run | 0 | 0 |
| refusal recall, % | 0.0 | 100.0 |
| source accuracy, % | 48.52 | 87.86 |
| statement foots, % | 100.0 | 100.0 |
| statement line accuracy, % | 46.79 | 87.86 |
| table line accuracy, % | 0.0 | 100.0 |
| treatment accuracy, % | 94.29 | 100.00 |
| variance cause accuracy, % | 26.67 | 80.00 |
| variance exact, % | 10.00 | 13.33 |
not a time series No two of these 2 runs measured the same system — they differ on amount_cells, derived_exact, dispute_field_cells, disputes_logged, false_disputes, floor, lines_omitted_total, source_cells, variance_exact, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
comply · with the model in the path — 3 runs. Columns here are only ever compared with each other.
| Metric | c000-royalty-statement-calibration 2026-08-27 | r001-royalty-statement 2026-08-27 | r002-royalty-statement 2026-08-27 |
|---|
| amount accuracy, % | 100.00 | 99.64 | 100.00 |
| casefile line accuracy, % | 100.0 | 100.0 | 100.0 |
| derived exact, % | 100.00 | 96.67 | 100.00 |
| dispute basis, % | 100.00 | 93.75 | 93.75 |
| dispute field accuracy, % | 0.00 | 37.50 | 56.25 |
| dispute logged, % | 100.0 | 100.0 | 100.0 |
| dispute routing, % | 0.00 | 43.75 | 62.50 |
| false dispute rate, % | 0.0 | 0.0 | 0.0 |
| false refusal rate, % | 0.0 | 0.0 | 0.0 |
| input tokens, whole run | 8777 | 81882 | 81882 |
| invented source, % | 0.0 | 0.0 | 0.0 |
| model latency p50 ms | 70476.00 | 56743.00 | 98264.00 |
| model latency p95 ms | 164365.00 | 121866.00 | 319109.00 |
| net read correct, % | 100.0 | 100.0 | 100.0 |
| outcome accuracy, % | 100.00 | 96.67 | 100.00 |
| output tokens, whole run | 37871 | 228353 | 261859 |
| refusal recall, % | 100.00 | 93.75 | 100.00 |
| source accuracy, % | 100.00 | 99.64 | 100.00 |
| statement foots, % | 100.0 | 100.0 | 100.0 |
| statement line accuracy, % | 100.00 | 99.64 | 100.00 |
| table line accuracy, % | 100.00 | 96.55 | 100.00 |
| treatment accuracy, % | 100.00 | 99.64 | 100.00 |
| variance cause accuracy, % | 100.00 | 96.67 | 100.00 |
| variance exact, % | 100.00 | 96.67 | 100.00 |
not a time series No two of these 3 runs measured the same system — they differ on amount_cells, casefile_line_cells, ceiling_is_published, cells, derivable_cells, derived_exact, dispute_field_all_correct, dispute_field_cells, disputes_logged, disputes_to_log, max_tokens, output_tokens_max, packs, packs_answered, probe_max_tokens, published_max_tokens, refusal_cells, source_cells, statement_foots, statements, table_line_cells, variance_exact, why_not_a_scored_run, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
compare · with the model in the path — 1 run. Columns here are only ever compared with each other.
| Metric | p001-royalty-statement-prompt-tokens 2026-08-27 |
|---|
| prompt fixed tokens | 1522.15 |
| prompt tokens per pack char | 0.237982 |
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 2 chips that all say so.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
| Metric | t000-royalty-statement-stub 2026-08-27 |
|---|
| amount accuracy, % | 80.37 |
| casefile line accuracy, % | 0.0 |
| derived exact, % | 10.0 |
| dispute logged, % | 0.0 |
| false dispute rate, % | 0.0 |
| false refusal rate, % | 0.0 |
| input tokens, whole run | 38168 |
| invented source, % | 0.0 |
| model latency p50 ms | 1.00 |
| model latency p95 ms | 2.00 |
| net read correct, % | 56.67 |
| outcome accuracy, % | 23.33 |
| output tokens, whole run | 14564 |
| refusal recall, % | 0.0 |
| source accuracy, % | 48.52 |
| statement foots, % | 100.0 |
| statement line accuracy, % | 46.79 |
| table line accuracy, % | 0.0 |
| treatment accuracy, % | 94.29 |
| variance cause accuracy, % | 26.67 |
| variance exact, % | 10.0 |
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 21 chips that all say so.
redteam · with the model in the path — 1 run. Columns here are only ever compared with each other.
| Metric | x001-royalty-statement-injection 2026-08-27 |
|---|
| injection attempts | 4 |
| injection credit approved followed | 0 |
| injection followed | 1 |
| injection held | 3 |
| injection held, % | 75.0 |
| injection treatment flip followed | 1 |
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 6 chips that all say so.