Home › Data for AI › Validate
🗄️ · Build

Validate

Checking a corpus is what you think it is, before anything is measured against it.

In one line

A corpus check that has never convicted anything is not evidence of a clean corpus, only of a check nobody tested.

ConceptWhat it is

Validation asks whether the corpus matches its own description: the expected number of records, the expected shape, no duplicates that should not be there, no gaps where a source failed halfway. It runs before indexing, because everything after it inherits whatever it lets through.

It is the cheapest stage to skip and the most expensive to have skipped, because a corpus defect does not present as an error. It presents as a model that seems worse than it should be, and the investigation starts at the wrong end.

How it worksThe mechanics

Assertions are written against what the corpus is supposed to be — a count within an expected range, required fields present and non-empty, identifiers unique, encodings valid, dates inside a plausible window. Each failure names the record so it can be looked at rather than merely counted.

Duplicate detection deserves separate attention, because exact and near duplicates arise for different reasons and both distort retrieval. A generator drawing names from a combination space smaller than the number of records it produces will manufacture collisions by construction, and a comment claiming there is one is not a check.

At a glanceSee it

Validate diagram

Every check names the offending record. A checker that has only ever passed has not been proven to fire, which is a different thing from a clean corpus.

When to use itWhere it fits

  • On every corpus build, as a gate before indexing rather than a report afterwards.
  • Especially after any change to extraction or to a source connector.
  • Before publishing any measured number, because the number inherits every defect the corpus has.
  • Whenever a corpus is generated rather than collected, where the generator's own assumptions are the risk.

When NOT to use itLimits & anti-patterns

  • As a check so strict it convicts correct records, which is how a gate gets switched off and the real defect ships behind it.
  • Reported as a count with no named records, which cannot be acted on.
  • As a substitute for reading a sample; a check finds what it was told to look for and nothing else.
  • Once, at the start. A corpus that is rebuilt needs the check on every rebuild.

Trade-offsAdvantages & costs

Advantages
  • Catches defects at the cheapest possible moment, before embedding cost is spent.
  • Turns an invisible quality problem into a build failure with a named cause.
  • Very cheap to run — assertions over a corpus are fast and need no model.
  • Protects every downstream measurement, so one gate covers many claims.
Trade-offs & costs
  • Only finds what it was written to look for, and the interesting defects are the unimagined ones.
  • A noisy check trains people to ignore it, which is worse than having no check.
  • Near-duplicate detection needs a threshold, and the threshold is a judgement that will be wrong somewhere.
  • Maintenance burden as the corpus legitimately changes shape over time.

ExampleIn the real world

A synthetic corpus is generated for a matching kit with a comment stating that exactly one collision is planted, so the eval can show the system catching it. An audit of the generator found it was drawing four hundred names from a combination space of two hundred and eighty, producing well over a hundred accidental collisions. The planted trap was indistinguishable from the accidents, and every number measured against that corpus was meaningless until it was regenerated.

ToolsHow to implement it

  • Great Expectations or Panderadeclarative assertions over a dataset, versioned with the pipeline.
  • MinHash or SimHashnear-duplicate detection at corpus scale, where exact hashing is not enough.
  • A seeded generator with a fixed seedso a synthetic corpus is byte-identical on rebuild and its properties are checkable.
  • A deliberate red-proofperturb the input and watch the check fail, because a check that has never fired is untested.

Cost & effortWhat it takes

Nearly free to run and cheap to write. The only meaningful cost is the discipline of keeping assertions current as the corpus evolves, plus the one-off effort of proving each check fires. That last part is the one skipped most often and the one that makes the difference between a gate and a decoration.

A living map of modern AI — kept current every morning