Home › Deployment, Inference & LLMOps › Experiment tracking
⚙️ · Operate

Experiment tracking

Recording what was tried, what changed and what it scored, so a re-run can be compared to the run it replaces.

In one line

Without tracking, a team relitigates the same three ideas every few weeks because nobody can show what happened the last time.

ConceptWhat it is

Experiment tracking is the record of attempts: for each run, the inputs that defined it — model version, prompt version, dataset version, parameters — and the results it produced, stored together and kept.

It is the least glamorous atom on this tile and the one that compounds fastest. LLM work is iterative by nature and most iterations do not help; the value is in knowing which ones did and, more usefully, which ones did not. A team without this record does not merely lose history, it repeats it, because a plausible idea that failed six weeks ago is still plausible today and there is nothing to consult.

How it worksThe mechanics

A run is a record with a stable identity, and the identity is qualified rather than generic — a name that says what was tried, not baseline, which collides the moment a second person uses it. The record carries every input version, the parameters, the results, and the prediction that was made before the run.

That last field is what makes the record worth reading afterwards. A run whose expected outcome was written down first tells you something whether it succeeds or fails; a run scored only in hindsight tends to be remembered as having confirmed whatever happened. The honest fields belong here too — what was not measured, what broke, what could not be verified — because a tracker holding only successes is a highlight reel and cannot support a decision about what to try next.

At a glanceSee it

Experiment tracking diagram

The prediction is written before the run, so the record says something whether it succeeds or fails. The failures are the half most often lost and most often needed.

When to use itWhere it fits

  • From the first comparison onwards, since the second run is already something the first should be compared against.
  • Whenever more than one person changes prompts or models, which is when memory stops being a shared resource.
  • Before any paid run, because an unrecorded paid run has to be bought again to be understood.
  • When a decision will be revisited — and prompt and model decisions always are.

When NOT to use itLimits & anti-patterns

  • As a platform decision; a directory of records with pinned inputs is a tracker.
  • Recording only the runs that worked, which produces a record that cannot warn anyone off anything.
  • Tracking runs whose inputs were not pinned — the record looks complete and the comparison is not valid.
  • As a substitute for a registry; tracking says what was tried, a registry says what is serving.

Trade-offsAdvantages & costs

Advantages
  • Stops the same idea being retried, which is the largest quiet cost in iterative work.
  • Makes a comparison valid by recording the inputs that have to match for it to mean anything.
  • Preserves negative results, which are the expensive half and the half that is normally lost.
  • Turns a paid run into a durable asset rather than a number somebody remembers.
Trade-offs & costs
  • Discipline decays unless the record is written by the run rather than by a person afterwards.
  • A record with unpinned inputs is worse than no record, because it invites an invalid comparison.
  • Volume makes finding the relevant prior run its own problem without good naming.
  • Tempts teams into re-scoring recorded outputs until the number improves, which is not a second run.

ExampleIn the real world

A team proposes raising a token ceiling to fix truncated outputs — an obvious diagnosis with an obvious fix. The tracker holds a run from a month earlier that did exactly that, and its honest fields record that the ceiling was not the constraint. A three-call probe confirms it again for a few cents, and the change that would have consumed a full run is not made. The value was not the successful run; it was the recorded failure that stopped the second one.

ToolsHow to implement it

  • A run record per attemptinputs, parameters, results, honest fields, one file.
  • MLflow or Weights & Biaseswhen comparison across many runs needs a real interface.
  • A qualified run idunique across the whole estate, because a generic id overwrites its neighbour in silence.
  • The prediction, written before the runthe field that makes a failed run informative.

Cost & effortWhat it takes

Effort is a convention and a file per run; the ceiling on usefulness is naming rather than storage. Its return is measured in runs not repeated, which on any kit that pays per call is the largest saving available — a recorded negative result costs nothing to consult and the run it prevents costs the full price.

A living map of modern AI — kept current every morning