Home › Data for AI › Transform
🗄️ · Build

Transform

Reshaping records into the unit the system retrieves, which quietly redefines what the eval measures.

In one line

Changing the shape of a record changes what a score means, so a transform is a change to the measurement and not only to the data.

ConceptWhat it is

Transformation reshapes cleaned records into the unit the system actually works with — splitting a long document into passages, joining a row to its context, pivoting an event stream into windows, collapsing a thread into a conversation.

It looks like plumbing and it is the stage that most often invalidates a comparison. If yesterday the unit was a page and today it is a section, then recall, precision and cost per question are all measured over different populations, and the two numbers are not comparable however similar they look.

How it worksThe mechanics

The unit is chosen deliberately and written down: what one retrievable record represents, and why that is the right granularity for the questions being asked. Splitting respects meaning where it can — a section boundary is a better split than a character count — and overlap is used where a passage would otherwise be cut mid-argument.

Because the unit defines the denominator of every metric, a change to it is versioned like a schema change. The eval set is re-derived or explicitly re-baselined, and the comparison across the change is either done properly or not claimed. A transform shipped quietly between two eval runs is how a system appears to improve without improving.

At a glanceSee it

Transform diagram

The unit is the denominator of every metric. Changing it without re-baselining is how a system appears to improve while nothing improved.

When to use itWhere it fits

  • Whenever the natural shape of the source is not the shape a question is asked about.
  • Splitting long documents, where whole-document retrieval would waste most of the context budget.
  • Joining a terse record to the context that makes it interpretable.
  • Windowing event streams, where a single record carries too little to be useful.

When NOT to use itLimits & anti-patterns

  • Between two eval runs whose numbers will be compared, unless the change is stated and the baseline re-derived.
  • Splitting purely by character count when structural boundaries are available and better.
  • Aggressively enough that a passage loses the context needed to interpret it.
  • As an unversioned change, since it silently redefines every metric downstream.

Trade-offsAdvantages & costs

Advantages
  • The right unit improves retrieval quality more reliably than most model or prompt changes.
  • Controls cost directly, because the unit determines how much context each answer consumes.
  • Structural splitting produces chunks that read as coherent passages, which is the practical test.
  • Making the unit explicit gives the team a shared vocabulary for a decision usually left implicit.
Trade-offs & costs
  • Silently invalidates comparisons across the change, and nothing warns you.
  • Re-transforming means re-embedding, which is the expensive part of a rebuild.
  • The best unit differs by question type, so a single choice is always a compromise.
  • Overlap increases index size and cost without improving every query.

ExampleIn the real world

A retrieval system moves from fixed thousand-character chunks to section-based splitting. Recall improves by a visible margin in the next eval run. It genuinely was better, but part of the gain was that the section unit produced fewer, richer records, so the denominator changed too. Nobody could say how much of the improvement was real until the run was repeated with the baseline re-derived.

ToolsHow to implement it

  • Structure-aware splitterssplitting on headings and sections rather than on a character count.
  • A written unit definitionone sentence stating what a record represents, versioned with the pipeline.
  • Overlap tuned by measurementchosen from eval results rather than from a default someone copied.
  • A corpus version bumped on every unit changeso a stale comparison is detectable rather than plausible.

Cost & effortWhat it takes

The transform itself is cheap; re-embedding after it is not, and that dominates the cost of any change here. Effort is low to implement and high to decide well, which is why the unit deserves an explicit decision rather than a default. Budget an eval re-baseline as part of any unit change.

A living map of modern AI — kept current every morning