Systems that improve are the ones where looking at failures is a scheduled job rather than something done after a complaint.
ConceptWhat it is
The improvement loop is the cycle that takes a system from working to getting better: capture what happened in production, read the failures, turn the recurring ones into an eval set, change something, and measure whether the change helped.
It is not a tool and cannot be bought. Every part of it is available in isolation — tracing, eval frameworks, judges — and teams routinely have all of them while the loop is not running, because the step that never gets scheduled is a person reading a hundred failures and naming what went wrong.
How it worksThe mechanics
Traces from production are sampled, weighted toward the ones a signal marks as suspect — a thumbs-down, a retry, an abandoned session, a guardrail that fired. Someone reads them and applies a label describing the failure, inventing categories as they appear rather than from a taxonomy decided in advance. Categories that recur become the axes you measure on.
Each recurring failure contributes cases to a held-out eval set with an expected outcome. A change — to the prompt, the retrieval, the chunking, the model — is then measured against that set rather than against impressions. The rule that keeps the loop honest is that the eval set is written down before the fix, so the fix cannot be tuned to the observed failure and then declared a success.
At a glanceSee it
The eval set is written before the fix. That ordering is what stops a change being tuned to the failure it was meant to generalise past.
When to use itWhere it fits
- As soon as a system is in front of real users, since real failures are the only ones worth prioritising.
- When quality complaints are anecdotal and there is no agreed measure to argue against.
- Before optimising cost, because a cheaper model is only cheaper if you can tell whether it got worse.
- Whenever a change is proposed on intuition; the loop is what converts an opinion into a measurement.
When NOT to use itLimits & anti-patterns
- Before there is traffic, where the honest first step is a small hand-built eval set instead.
- As a reason to delay shipping — the loop needs production to feed it.
- As a purchased capability; buying a tracing tool is not the same as scheduling the reading.
- Where the eval set is allowed to change with each fix, at which point the number stops meaning anything.
Trade-offsAdvantages & costs
Advantages
- Converts vague quality debate into a measurement everyone can check.
- Prioritises by observed frequency, so effort goes to what actually fails rather than what is feared.
- Catches regressions from model, prompt and data changes that nobody would think to test.
- The eval set compounds — every cycle leaves the team permanently better able to judge the next change.
Trade-offs & costs
- The reading step is genuine human effort and cannot be delegated to a model without losing what makes it work.
- An eval set drifts from production over time and needs refreshing, which is its own recurring cost.
- Judge-based scoring introduces a second thing that can be wrong, and it needs validating against human labels.
- It only measures what the set covers, so a category nobody labelled stays invisible however many cycles run.
ExampleIn the real world
A document assistant is described as unreliable, with no agreement on why. Two hundred sampled traces are read and labelled in an afternoon. Four categories account for most of it, and the largest is not a model failure at all: tables extracted from PDFs arrive as run-on text, so the model answers confidently from mangled input. The fix is in the extractor, and it would never have been proposed while the conversation was about prompts.
ToolsHow to implement it
- Langfuse or Phoenixtrace capture with the retrieval detail intact, which is what makes a trace readable.
- A spreadsheetgenuinely the right tool for the labelling pass; the value is in reading, not in tooling.
- promptfoo, Braintrust or a plain test harnessrunning the eval set repeatably against each candidate change.
- An LLM judge validated against human labelsuseful for scale, but only after you have measured how far it agrees with you.
Cost & effortWhat it takes
The measurable cost is small — sampling and re-running an eval set of a few hundred cases is cents to low dollars per cycle. The real cost is a person's attention for a few hours per cycle, which is the part teams try to skip and the part that produces the finding. Effort is low to start and the discipline is what is hard to sustain.
What changedWhat changed here
- An Anthropic researcher just gave us a peek at self-improving AI
An Anthropic researcher showed automated systems improving performance on all 10 misaligned-behavior benchmarks without degrading overall performance. That's a concrete early signal that self-improvement for safety is moving from theory to a working loop — worth watching for how model-training economics change.
Three kinds of claim, strongest first. Signal runs every morning.