Evals tell you what a system did once; alerting is what makes them tell you when that stops being true.
ConceptWhat it is
Alerting on drift is the step that closes the gap between an eval suite and production. An eval run establishes a baseline — this system, on this set, behaved this way. Alerting watches live traffic for deviation from that baseline and says so while there is still time to act.
It matters because nothing else in the stack notices this class of failure. An LLM system that has drifted does not throw. Latency is fine, error rates are fine, every dashboard is green, and the answers are worse. Traditional monitoring watches whether the system is running; drift alerting watches whether it is still right, and those are different questions with different signals.
This is also where evals stop being a gate and become a control. A suite that runs only before release protects the release; the same suite sampled continuously against live traffic protects the service, and the deviation it reports is the input a guardrail acts on.
How it worksThe mechanics
Three things drift and they are worth separating, because they have different causes and different fixes. Input drift is the traffic changing shape — new topics, longer documents, a different language mix. Output drift is the responses changing without the inputs having changed, which usually means the model or a prompt moved. Quality drift is the measured score falling, and it is the one that actually matters; the other two are early warnings of it.
The measurement is a continuous sample rather than a full re-run. A small slice of live traffic is scored on the same rubric the offline suite uses, producing a rolling number comparable to the baseline. Cheap proxies run on everything — refusal rate, response length distribution, retrieval hit rate, the share of answers citing no source — because they cost nothing and move early.
Thresholds are set from the baseline's own variance rather than from a round number, which is the difference between an alert that gets read and one that gets muted. Each alert names what deviated, by how much, and against which baseline. And the guardrail response is decided in advance: log it, route affected traffic to a fallback, or hold the release — because an alert with no prepared response is a notification.
At a glanceSee it
The eval run sets the baseline and its variance sets the thresholds. A continuous sample and cheap proxies both feed them, and every alert has a response decided before it fires.
When to use itWhere it fits
- Once a system is serving real traffic, which is the point at which offline evals stop being sufficient on their own.
- Wherever the corpus, the model or the prompts can change without a deploy — all three are common and none is visible to uptime monitoring.
- When a quality failure is expensive and silent, which is the normal shape of LLM failure.
- Alongside a guardrail that can actually act, since detection without a prepared response only shortens the surprise.
When NOT to use itLimits & anti-patterns
- Before a baseline exists; an alert with nothing to deviate from is a threshold somebody invented.
- On a metric too noisy to hold a threshold, where the result is alert fatigue and then a muted channel.
- Scoring every request with a judge model, which is a cost curve that grows with traffic for very little extra signal over a sample.
- As a replacement for offline evals — drift alerting tells you something changed, not what to do about it.
Trade-offsAdvantages & costs
Advantages
- Catches the failure class nothing else in the stack can see, because the system stays healthy while getting worse.
- Turns an eval suite from a release gate into a live control, which is most of its unrealised value.
- Localises a regression by which signal moved first — input, output or score.
- Feeds the improvement loop with real failures rather than with reported ones.
Trade-offs & costs
- Thresholds are genuinely hard to set and a bad one is worse than none, because a muted alert is indistinguishable from a healthy system.
- Continuous scoring costs money on every sampled call, so the sample rate is a real budget decision.
- A judge model used as the scorer can drift itself, which needs its own pinned version.
- Input drift fires on legitimate change as readily as on problems, so it needs a human reading it rather than an automatic action.
ExampleIn the real world
A support assistant's answers get worse over about ten days with no deploy, no error rate change and no latency change. The proxy that moves first is the share of answers citing no retrieved source, up from four per cent to nineteen. The sampled quality score follows it down a few days later. The cause is upstream: a documentation migration had begun publishing pages in a template the extractor handled badly, so retrieval was returning less and the model was answering from its own weights. Uptime monitoring had nothing to say about any of it.
ToolsHow to implement it
- The offline eval suite, sampled continuouslythe same rubric against live traffic is what makes the two numbers comparable.
- Cheap proxy metricsrefusal rate, length distribution, retrieval hit rate, share of answers citing no source.
- Baseline variancethresholds derived from how much the metric already moves, not from a round number.
- A prepared guardrail responselog, route to a fallback, or hold the release, decided before the alert fires.
Cost & effortWhat it takes
The proxies are effectively free — they are counts over data already being logged. The sampled scoring is the real cost and it scales with the sample rate rather than with traffic, which is what makes it affordable: a small slice scored well beats everything scored cheaply. Budget it as a running line rather than as a project, and set the rate from how fast you need to notice rather than from how much you can afford to score.