Home › Deployment, Inference & LLMOps › Dataset versioning
⚙️ · Operate

Dataset versioning

Giving a corpus an immutable identity so a result computed against it can be reproduced rather than re-argued.

In one line

A dataset without a version is a moving target, and every number measured against a moving target expires quietly.

ConceptWhat it is

Dataset versioning means a corpus has an identity that will still resolve to the same bytes next year. Not a folder name, not a date in a filename, but a reference that either returns exactly what it returned before or fails loudly.

The practice earns its own page because it is the assumption every eval quietly makes. Comparing this week's score to last week's is only meaningful if the set is the same set, and the failure mode is not an error — it is a comparison that still renders, still looks like a trend, and is measuring two different things.

How it worksThe mechanics

The corpus is written once and never edited in place. A new version is a new write with a new identifier, and the old one stays retrievable. Content addressing is the strongest form of this: the identifier is a hash of the content, so an identifier that resolves at all is proof the bytes are unchanged.

The version then has to reach the places that cite it. An eval record carries the dataset version alongside the model and prompt versions, so a stored result names all three of its inputs. Held-out sets get the same treatment and one extra rule — adding examples creates a new version rather than growing the existing one, because a set that grows silently makes every historical score incomparable.

At a glanceSee it

Dataset versioning diagram

A corpus written once and cited by version is what makes two scores comparable. A set that grows in place makes every historical number quietly incomparable.

When to use itWhere it fits

  • For every held-out eval set, without exception, because that is the set whose stability the whole practice depends on.
  • Whenever a corpus refreshes on a schedule and results are compared across refreshes.
  • When more than one person can add to a set, since that is when it starts moving without announcement.
  • Before publishing any number that someone might reasonably ask you to reproduce.

When NOT to use itLimits & anti-patterns

  • As a reason to adopt a versioning platform; a write-once bucket and a hash go a long way.
  • Versioning derived artefacts that can be rebuilt deterministically from a versioned source — version the source instead.
  • For scratch data during exploration, where the ceremony outweighs the benefit.
  • As a substitute for lineage; a version says which bytes, lineage says where they came from and what used them.

Trade-offsAdvantages & costs

Advantages
  • Makes a result reproducible, which is the property that separates a measurement from an impression.
  • Turns a suspected regression into a controlled comparison by holding the data constant.
  • Makes set growth visible instead of silent, which is where most incomparable numbers come from.
  • Cheap for text corpora, where storage is small relative to almost everything else in the system.
Trade-offs & costs
  • Storage grows monotonically, which needs a retention policy rather than a cleanup habit.
  • Discipline is the real cost — a single in-place edit undoes the guarantee for every version after it.
  • Large binary corpora make the write-once rule genuinely expensive.
  • A version id is useless unless the code that consumes the set is required to record it.

ExampleIn the real world

A team reports steady improvement on their eval set over six weeks, then cannot reproduce any of it. The set had been appended to twice during the period by different people, each addition drawn from recent production traffic and therefore easier than the original examples. Every score in the trend was real and every comparison between them was meaningless, because none of the six numbers had been computed against the same set.

ToolsHow to implement it

  • Content-addressed storagethe identifier is a hash, so resolving it proves the bytes are unchanged.
  • DVC or LakeFSversion control semantics over a bucket, when a team needs branches and tags rather than ids.
  • A write-once bucket with object lockthe cheapest form that actually holds.
  • The version field on every eval recordwithout it the versioning exists and nothing cites it.

Cost & effortWhat it takes

For text corpora the storage cost is close to noise and the effort is a convention plus one field. It rises sharply for image or audio sets, where write-once means genuinely duplicating large objects and a retention policy stops being optional. The cost of not doing it is paid once, all at once, when a published number has to be defended and cannot be recomputed.

A living map of modern AI — kept current every morning