HistoryRun history
9 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 4 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 4 runs. Columns here are only ever compared with each other.
| Metric | a001-deviation-intake-norubric 2026-08-25 | a002-deviation-intake-blind 2026-08-25 | c000-deviation-intake-calibration 2026-08-25 | r001-deviation-intake 2026-08-25 |
|---|
| basis accuracy, % | 84.0 | 86.0 | 100.0 | 94.0 |
| class accuracy aligned register, % | 94.12 | 94.12 | 100.00 | 100.00 |
| class accuracy header blank, % | 83.33 | 91.67 | 100.00 | 91.67 |
| class accuracy header complete, % | 90.0 | 100.0 | 100.0 | 100.0 |
| class accuracy header wrong, % | 75.0 | 75.0 | 100.0 | 100.0 |
| class accuracy misleading register, % | 68.75 | 93.75 | 100.00 | 93.75 |
| critical recall, % | 64.29 | 92.86 | 100.00 | 92.86 |
| deviation class accuracy, % | 86.0 | 94.0 | 100.0 | 98.0 |
| distributed accuracy, % | 100.0 | 96.0 | 100.0 | 98.0 |
| evidence quote present, % | 100.0 | 100.0 | 100.0 | 100.0 |
| evidence quote verbatim, % | 100.0 | 100.0 | 100.0 | 100.0 |
| false critical, % | 2.78 | 5.56 | — | 0.00 |
| input tokens, whole run | 49760 | 56560 | 7530 | 61860 |
| model latency p50 ms | 6624.00 | 6491.00 | 6697.00 | 4548.00 |
| model latency p95 ms | 44502.00 | 47834.00 | 47254.00 | 30432.00 |
| output tokens, whole run | 64104 | 74065 | 7725 | 50320 |
not a time series No two of these 4 runs measured the same system — they differ on basis_correct, basis_of, class_accuracy_alarming_but_minor_correct, class_accuracy_alarming_but_minor_of, class_accuracy_aligned_register_correct, class_accuracy_aligned_register_of, class_accuracy_flat_but_critical_correct, class_accuracy_flat_but_critical_of, class_accuracy_header_blank_correct, class_accuracy_header_blank_of, class_accuracy_header_complete_correct, class_accuracy_header_complete_of, class_accuracy_header_wrong_correct, class_accuracy_header_wrong_of, class_accuracy_misleading_register_correct, class_accuracy_misleading_register_of, class_accuracy_of, class_accuracy_truth_critical_correct, class_accuracy_truth_critical_of, class_accuracy_truth_major_correct, class_accuracy_truth_major_of, class_accuracy_truth_minor_correct, class_accuracy_truth_minor_of, class_accuracy_truth_not_critical_correct, class_accuracy_truth_not_critical_of, critical_caught, critical_missed, critical_of, deviation_class_correct, deviation_class_of, distributed_correct, distributed_of, documents, evidence_quote_present, evidence_quote_present_of, evidence_quote_verbatim, evidence_quote_verbatim_of, false_critical, false_critical_of, over_called, regulator_notified_correct, regulator_notified_of, under_called — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
| Metric | b000-deviation-intake-prior 2026-08-25 | b001-deviation-intake-lexicon 2026-08-25 | b002-deviation-intake-fieldrules 2026-08-25 |
|---|
| basis accuracy, % | 36.0 | 46.0 | 72.0 |
| class accuracy aligned register, % | 52.94 | 82.35 | 82.35 |
| class accuracy header blank, % | 33.33 | 50.00 | 58.33 |
| class accuracy header complete, % | 36.67 | 60.00 | 100.00 |
| class accuracy header wrong, % | 37.5 | 50.0 | 37.5 |
| class accuracy misleading register, % | 0.0 | 0.0 | 75.0 |
| critical recall, % | 0.00 | 21.43 | 78.57 |
| deviation class accuracy, % | 36.0 | 56.0 | 80.0 |
| distributed accuracy, % | 80.0 | 76.0 | 92.0 |
| evidence quote present, % | 0.0 | 0.0 | 0.0 |
| false critical, % | 0.00 | 13.89 | 5.56 |
| input tokens, whole run | 0 | 0 | 0 |
| model latency p50 ms | 0.00 | 0.00 | 0.00 |
| model latency p95 ms | 0.00 | 0.00 | 0.00 |
| output tokens, whole run | 0 | 0 | 0 |
not a time series No two of these 3 runs measured the same system — they differ on basis_correct, class_accuracy_alarming_but_minor_correct, class_accuracy_aligned_register_correct, class_accuracy_flat_but_critical_correct, class_accuracy_header_blank_correct, class_accuracy_header_complete_correct, class_accuracy_header_wrong_correct, class_accuracy_misleading_register_correct, class_accuracy_truth_critical_correct, class_accuracy_truth_major_correct, class_accuracy_truth_minor_correct, class_accuracy_truth_not_critical_correct, critical_caught, critical_missed, deviation_class_correct, distributed_correct, false_critical, over_called, regulator_notified_correct, under_called, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
| Metric | t000-deviation-intake-stub 2026-08-25 |
|---|
| basis accuracy, % | 36.0 |
| class accuracy aligned register, % | 52.94 |
| class accuracy header blank, % | 33.33 |
| class accuracy header complete, % | 36.67 |
| class accuracy header wrong, % | 37.5 |
| class accuracy misleading register, % | 0.0 |
| critical recall, % | 0.0 |
| deviation class accuracy, % | 36.0 |
| distributed accuracy, % | 80.0 |
| evidence quote present, % | 0.0 |
| false critical, % | 0.0 |
| input tokens, whole run | 23331 |
| model latency p50 ms | 0.00 |
| model latency p95 ms | 0.00 |
| output tokens, whole run | 1750 |
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 15 chips that all say so.
redteam · with the model in the path — 1 run. Columns here are only ever compared with each other.
| Metric | x001-deviation-intake-injection 2026-08-25 |
|---|
| any field change rate, % | 35.71 |
| any field changed | 5 |
| suppressed | 0 |
| suppression rate, % | 0.0 |
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 4 chips that all say so.