COUNTED247 · 104 · 192 · 212 / 270afe memo completeness pct — AFE memo completeness -- disposition, finding, governing term, support, and on a miscoded charge the well it belongs toDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED247 · 137 · 192 · 212 / 270disposition accuracy pct — disposition only, three-way -- what a flagger can scoreDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED115 · 33 · 96 · 96 / 120findings caught pct — flaggable lines caughtDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED13 · 26 · 34 · 34 / 130overflag rate pct — standing lines OVER-FLAGGEDDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED92 · 33 · 96 · 96 / 96structured finding caught pct — structured findings caught -- the review's own tables prove themDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED23 · 0 · 0 · 0 / 24prose finding caught pct — PROSE findings caught -- only a cost review note reveals themDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED13 · 26 · 34 · 34 / 34silent overflag pct — SILENT OVER-FLAG -- lines whose apparent breach the review itself already answersDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED16 · 0 · 0 · 20 / 20insufficient caught pct — unsettleable lines named as unsettleableDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED115 · 24 · 96 · 96 / 115finding accuracy pct — finding named correctly, on lines flaggedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED115 · 0 · 96 · 96 / 115term accuracy pct — governing term cited correctlyDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED115 · 0 · 96 · 96 / 115support completeness pct — required support attachedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED200 · 0 · 0 / 257fabricated citation pct — lines citing an id the review never printsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED6 · 0 · 0 · 0 / 270silent omission pct — lines left off the memo entirelyDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED37 · 28 · 37 · 38 / 45afe position accuracy pct — AFE position, three-wayDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED23 · 0 · 0 · 0 / 24misattribution named of all pct — WHICH WELL a miscoded charge belonged to, over every miscoded line in the corpusDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
How the method was validated
evals/check_labels.py re-derives every structured finding from the reviews themselves and asserts ELEVEN properties over 270 cost lines -- that every gold line is in its review with the same category and source document, that every term and support identifier the key cites is PRINTED there, that a structured FLAG re-derives to the same finding, term and support, that a clean line re-derives to WITHIN_AUTHORITY, that an unsettleable line re-derives to INSUFFICIENT_EVIDENCE, that a prose FLAG looks CLEAN to the tables while a prose trap looks FLAGGED to them with the finding its case implies, that the AFE position follows from the line dispositions, that the withheld section never reaches the prompt, that every case agrees with its own disposition, finding and channel, that no category appears twice in one review, and that no belongs_to names a well the review does not print. 0 violations. Red-proven by seeding five defects one at a time and convicting all five by name.