Home › Operate
🚨 Debugging guide Operate

Failure Modes

How AI systems break in production, what each failure looks like, and what fixes it.

At a glanceFailure Modes

Failure Modes diagram

Triage order. Every question splits the system in half before it splits again — the fastest bug is the one you localise, not the one you theorise about.

Why it mattersKnowing how it works is not knowing how it breaks

Everything else on this site teaches you to build. This page teaches you to debug — and that is the difference between someone who has demoed AI and someone who has shipped it. Production AI fails differently from production software. There is no stack trace. The system does not crash; it returns a fluent, confident, well-formatted, wrong answer, with a 200 status code, in 900ms, and your monitoring stays green.

So the discipline is different. You are not looking for an exception — you are looking for the stage that went wrong in a pipeline where every stage produces plausible output. This page is organised the way you will actually use it: symptom → likely cause → what to check first.

Rule

Read the trace, not the answer. The answer tells you the system failed; only the trace tells you where.

The methodBisect the pipeline before you theorise

A RAG or agent system is a chain: query → retrieve → assemble context → generate → post-process. Every debugging session is the same move — cut the chain in half and ask which side the failure is on. The single highest-value question in applied AI:

Was the right information in the context window when the model generated?

  • No→ you have a retrieval bug. The model never had a chance. Stop tuning the prompt — you are prompting a model that cannot see the answer.
  • Yes, and it still got it wrong→ you have a generation bug. Stop tuning the retriever. More chunks will make it worse, not better.
  • Yes, and the output is what you literally asked for→ you have a specification bug. The system is fine; your definition of correct is wrong.

Teams routinely spend a week on the wrong half of that fork. The check costs 30 seconds: log the retrieved chunks, and read them yourself before you read the answer. If the answer is not in what you retrieved, nothing downstream matters.

The academic version of this fork: Barnett et al. engineered three production RAG systems and catalogued seven failure points, of which three are retrieval-side and four are generation-side. Their conclusion is the uncomfortable one — "validation of a RAG system is only feasible during operation," and robustness "evolves rather than [is] designed in at the start."

Symptom 1RAG returns garbage

The most common production complaint, and the most commonly misdiagnosed — because "the answer is wrong" is a symptom of at least seven distinct diseases. Barnett et al. name them; this is the map from their failure points to what you actually do about it.

Failure pointWhat you observeHalfFirst thing to check
FP1 Missing contentConfident answer to a question the corpus cannot answerCorpusIs the document even ingested? Grep the raw store, not the index.
FP2 Missed top-rankedAnswer exists, was retrieved, ranked #23 of top-20RetrievalRaise K and re-check. If it appears, you have a ranking problem — add a reranker.
FP3 Not in contextRetrieved but dropped during consolidationAssemblyYour token budget or dedupe step is silently truncating. Log context length per call.
FP4 Not extractedAnswer is in the context; model missed itGenerationNoise and contradiction in context. Retrieve fewer, better chunks.
FP5 Wrong formatProse when you asked for a tableGenerationInstruction lost in a long prompt. Move format spec to the end; use structured output.
FP6 Incorrect specificityToo vague or too narrow to be usefulSpecQuery/answer granularity mismatch. Usually a chunk-size decision, not a prompt one.
FP7 IncompleteCorrect but missing half the available answerGenerationMulti-chunk synthesis. Model anchored on chunk 1. Check ordering and position.

Failure points from Barnett et al., Seven Failure Points When Engineering a Retrieval Augmented Generation System (2024), drawn from research, education and biomedical case studies. The "half" and "first check" columns are the operational reading.

Chunk boundaries — where meaning gets amputated

Chunking is where most RAG quality is silently lost, because the failure is invisible at index time and only shows up as bad answers months later. A chunk is embedded as if it were a standalone document. It usually is not. Split a 10-K at a fixed 500 tokens and you get a chunk reading "revenue grew 3% over the prior period" — no company, no year, no segment. It is lexically fluent and semantically orphaned. It will never be retrieved for "How did ACME grow in Q2 2025?" because it does not contain "ACME", "grow" is a weak signal, and the embedding has nothing to anchor to.

Anthropic's Contextual Retrieval attacks exactly this: prepend a short, LLM-generated description situating each chunk in its parent document before embedding it. The measured effect on top-20 retrieval failure rate:

SetupTop-20 retrieval failure rateReduction
Baseline embeddings5.7%—
+ Contextual embeddings3.7%35%
+ Contextual embeddings & contextual BM252.9%49%
+ Reranking1.9%67%

Source: Anthropic, Introducing Contextual Retrieval. Note the shape of the result — hybrid search (dense + BM25) and reranking are doing enormous work. If you are running pure dense retrieval with no reranker, that is the cheapest fix available to you, and you should do it before anything more exotic.

Symptom: retrieval works for keyword-ish queries and collapses on conceptual ones. Cause: you are relying on lexical overlap you did not know you had. Check: take a failing query, search the corpus for its literal keywords. If that works and the vector search does not, your embeddings are not carrying the semantics you assumed.

Symptom: the answer is split across a boundary and you always get half. Cause: fixed-size chunking cut a table, list, or argument in two. Check: print the chunk. Look at where it starts and ends. Chunk on structure (headings, sections) with overlap, not on a token count you picked because it was round.

Embedding mismatch and the query/document asymmetry

Two failures that look identical from the outside and have completely different fixes.

Embedding mismatch is the mechanical one: the vectors in your index were produced by a different model — or a different version, or a different pooling/normalisation config — than the vectors you are querying with. The two live in incompatible spaces, so cosine similarity returns something, but it means nothing. The tell is diagnostic gold: results are not just bad, they are uncorrelated — nonsense with confident scores, and the same nonsense for every query. Ordinary bad retrieval degrades gracefully and stays topically adjacent; a space mismatch does not. Check the model ID and dimension stamped on the index against the one your query path loads. If your index has no such stamp, that is the bug — fix that first. Re-embedding the corpus is the only fix, which is why changing embedding model is a migration, not a config change.

Query/document asymmetry is the subtle one. Sentence-Transformers draws the distinction explicitly: symmetric search has query and corpus entries of similar length and kind (finding duplicate questions — you could swap them); asymmetric search has a short query hunting a long passage ("What is Python" → a paragraph). Their guidance is blunt — it is "critical that you choose the right model," and a model trained for symmetric similarity applied to asymmetric retrieval is miscalibrated for exactly the length and content mismatch your product is built on. Nearly every RAG system is asymmetric. A meaningful number are built on symmetric-trained embeddings because that was the default in the tutorial.

The asymmetry is real even with the right model: a question and its answer do not look alike. "Why did the deployment fail?" shares almost no surface with a log excerpt that explains it. This is what HyDE (Gao et al.) exploits — generate a hypothetical answer to the query, embed that, and search with it, so you are matching document-shaped text against documents. The encoder's "dense bottleneck filters out the incorrect details" of the invented answer while keeping its relevance structure. You are trading a model call for retrieval quality — the right trade when queries are short and vague, the wrong one when latency is the constraint.

Lost in the middle — and why more context is not more knowledge

The most expensive false intuition in RAG: the context window is 200K, so let us just put everything in it. Liu et al. showed models do not read the context window uniformly. Performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts" — and, critically, this holds even for models explicitly built for long context. Your top-ranked chunk sitting at position 12 of 20 is in the worst seat in the house.

It gets worse when you remove the crutch. Needle-in-a-haystack tests mostly pass because the needle and the question share literal words — the model can pattern-match without understanding. NoLiMa rebuilt the test with minimal lexical overlap, forcing genuine latent association. Across 13 models all claiming 128K+ context: strong below 1K, and at 32K, 11 of the 13 fell below 50% of their own short-context baseline. GPT-4o, one of the best performers, dropped from 99.3% to 69.7%. The authors attribute it to the attention mechanism struggling in long contexts when literal matches are absent — and note that reasoning models and CoT prompting do not rescue it.

So

A 128K context window is a capacity spec, not a comprehension guarantee. Retrieval precision is not obsolete — long context made it invisible, which is worse. Put your best chunk first, your second-best last, and stop padding the middle with hopeful filler.

Symptom 2Hallucination — a taxonomy, because the kinds have different fixes

"The model hallucinates" is not a bug report; it is a genre. Treating all fabrication as one problem is why teams throw RAG at it, watch it not go away, and conclude the model is bad. Ji et al.'s survey gives the load-bearing split:

KindDefinitionExampleFix
IntrinsicOutput that contradicts the source contentSource says 2019; model says 2021Groundedness check against the retrieved context. Cheap and mechanical — you have the source, so you can verify.
ExtrinsicOutput that cannot be verified from the source — neither supported nor contradictedA plausible extra clause nobody provided; possibly even trueHard. Requires an external truth source or a citation requirement. Cannot be caught by comparing to context alone.

Definitions quoted from Ji et al., Survey of Hallucination in Natural Language Generation. The distinction matters commercially: intrinsic hallucination is a solvable engineering problem, extrinsic hallucination is a product-design problem about what you are willing to claim.

The survey also separates the causes, which is where fixes actually come from:

Why it fabricates

  • Source-reference divergencethe training targets themselves contained information not supported by their sources. The model learned that inventing is what this task looks like. Relevant to you if you fine-tune: your data teaches the behaviour you are complaining about.
  • Parametric knowledge biasthe model prefers what it memorised over what you gave it. This is the RAG killer: you retrieved the correct, updated document and the model answered from pre-training anyway. Symptom: the answer is plausibly out of date rather than random. Fix: instruct explicitly to answer only from context, demand citations, and eval for it.
  • Erroneous decodingand imperfect representation learning — attention and sampling errors. Where temperature and decoding strategy actually help. Note this is the smallest of the causes and the one teams reach for first.
  • Exposure biastraining on ground truth, inference on its own output. Errors compound along a long generation. Symptom: the answer starts fine and degrades. This is also why agent chains rot.

And then the structural cause. Kalai et al. (OpenAI) argue models hallucinate because training and evaluation reward guessing over admitting uncertainty — like a student on an exam where blank scores zero and a guess might score one. Benchmarks "overwhelmingly penalize uncertainty," so a model that says "I don't know" loses to one that bluffs. Their proposed fix is socio-technical: penalise confident errors more than abstentions.

The operational consequence is immediate and under-appreciated. If your own eval marks "I don't know" as a failure, you are actively training your product to lie to your users. Score abstention separately from error. An honest "not in the documents" is a correct answer to an unanswerable question — and per FP1 above, your corpus is full of them.

Symptom 3Agents loop, stall, or confidently do the wrong thing

Agents fail in ways that look like bugs and are usually design. The best evidence base here is Cemri et al., who annotated 1,600+ execution traces across 7 multi-agent frameworks and derived MAST — 14 failure modes in 3 categories, at 0.88 inter-annotator agreement. The headline finding is the one worth internalising: failures cluster in system design, inter-agent misalignment, and task verification — that is, in specification and coordination, not in model intelligence. A smarter model does not fix a system whose agents disagree about what they are doing.

SymptomWhat it usually meansCheck first
Loops the same tool callTool returned something the agent cannot interpret as success — often an empty result or an error string it reads as "try again". No progress signal exists.The tool's return value on the failing path. Empty list vs error vs null. Make failure legible: return "no results, do not retry" not [].
Ping-pongs between two stepsTwo subgoals whose success conditions contradict. Classic inter-agent misalignment.Read the trace for the point where the goal restates. Cap iterations and log the termination reason — always.
Stalls / does nothingAmbiguous termination condition; the agent believes it is done, or cannot select among equal options.Is "done" defined as a checkable predicate, or as vibes in a prompt?
Confidently wrong, one clean passNo verification step. Nothing checked the work. The agent is not lying — nobody asked.Whether any step can return "this is wrong, redo it". Most agent systems have zero verifier.
Works solo, fails in the crewContext loss at handoff. Agent B did not receive what Agent A knew.Log exactly what crosses the boundary. It is almost always less than you think.
Degrades as the run gets longerExposure bias + context growth. Early errors compound and the middle of a long history is not read (Liu et al.).Trace length at failure. Summarise or checkpoint state instead of appending forever.

Symptom taxonomy informed by Cemri et al., Why Do Multi-Agent LLM Systems Fail? (MAST); the compounding-over-length mechanism connects to exposure bias (Ji et al.) and lost-in-the-middle (Liu et al.).

The three controls that prevent most agent incidents. 1) A step cap with a logged termination reason — an agent that stops at 10 steps and says why is debuggable; one that stops at 10 steps silently is not, and one with no cap is an open invoice. 2) A verifier — some step that can fail the work. MAST's third category is literally task verification. 3) Legible tool errors — the agent can only reason about what your tools tell it, and most tools tell it nothing useful.

Symptom 4Quality dropped and nothing changed

The most disorienting failure, because your diff is empty. Something did change — just not in your repo. In rough order of how often it is the culprit:

CauseTellCheck
Data driftGradual decay over weeks; failures concentrate in new topicsCompare the query distribution this month vs launch. Users found a use case you never indexed.
Cache stalenessConfidently outdated answers; correct on cache-miss pathsBypass the cache and re-run the failing query. If it is right, your cache is the bug, and the index may be stale relative to the source of truth.
Prompt driftSlow degradation nobody can datePrompt history. Six people appended six edge-case clauses; the format instruction is now buried at position 400 of 900 — see lost-in-the-middle. Prompts need review and version control like code.
Silent version bumpStep change on a specific date; no deployAre you pinned to a dated snapshot or an alias? An unpinned alias means the vendor deploys to your production.
Model deprecationHard errors, or silent redirect to a successorThe vendor deprecation page. OpenAI commits to at least 6 months notice for GA models and at least 3 months for specialised variants — but notice arrives by email to whoever created the account, not to you.
Non-determinismSame input, different output, temperature 0Not a bug you can fix — see below. Fix the expectation, and the eval that assumed exact match.

The non-determinism one deserves its mechanism, because the folk explanation is wrong and the wrong explanation makes it unfixable. "GPU concurrency plus floating-point non-associativity" is not the cause: "even on a GPU, running the same matrix multiplication on the same data repeatedly will always provide bitwise equal results." The real cause is that inference kernels are not batch-invariant — their numerics shift with batch size — and batch size varies with server load, which is other people's traffic. So your output at temperature 0 depends on how busy the provider was. Two properties compose into the surprise: kernels lack batch invariance (mathematical) and load varies nondeterministically (operational).

Consequence

Temperature 0 is not a reproducibility guarantee. Any eval asserting exact string equality will flake, and you will waste a week debugging your code instead of your assumption.

The prophylactic: make "nothing changed" a checkable statement

  • Pin model versions to dated snapshots.An alias is a promise someone else can break.
  • Version prompts and record the prompt hash with every trace."Which prompt produced this?" must be answerable in one query.
  • Stamp the indexwith embedding model, dimension, and build date. Half of "nothing changed" is an index that was rebuilt by a cron job nobody remembers writing.
  • Subscribe to the deprecation feedfor every provider you use, to a channel a human reads.
  • Track the query distributionnot just latency and errors. Drift is visible there months before it is visible in complaints.

Symptom 5The eval passes and users still complain

Your dashboard is green. Support is not. One of these is measuring your product; it is not the dashboard.

The four eval/reality gaps, in order of frequency

  • You bought generic metrics instead of building specific ones.Hamel Husain's position, from working failed LLM products: unsuccessful products "almost always share a common root cause: a failure to create robust evaluation systems," and the trap is believing the right framework will supply them. Generic metrics are worse than useless — they manufacture a false sense of progress while teams track vanity numbers uncorrelated with real user problems. His alternative is unglamorous: error analysis — look at your data, categorise the errors you actually have, and write a test for each one you find. In a real CRM deployment the real bugs were things like the system prompt's UUID leaking into messages — which no off-the-shelf metric was ever going to catch.
  • Your eval set is stale.It was built at launch from queries you imagined. Users have since invented usage you did not predict. The eval passes because it is testing last year's product. Every production failure should end by appending a case to the eval set — that is what makes the set track reality rather than memory.
  • The judge is biased.LLM-as-judge is the standard tool and it works — Zheng et al. found strong judges reach over 80% agreement with human preferences, the same level humans agree with each other. But the same paper names the biases: position bias (favouring an answer by where it sits), verbosity bias (longer reads as better), and self-enhancement bias (favouring output from similar models). Verbosity bias is the dangerous one commercially: your judge rewards padding, so your prompts evolve toward padding, so your outputs get longer, so your bill grows — and your users, who wanted an answer, get an essay. The judge score rises the whole time. Mitigate by swapping positions and averaging, and by never letting the judge be the same model family as the generator without checking that assumption.
  • You are measuring the pipeline, not the outcome.Retrieval precision, faithfulness and latency can all pass while the product fails, because none of them ask "did the user get what they came for?" These are diagnostics — they tell you which half is broken when something is. They are not a definition of success.
Goodhart

Every eval metric becomes a target and then stops being a measure. The defence is not a better metric — it is looking at real traces every week, by hand, forever. There is no version of this job that does not involve reading what your system actually said.

Symptom 6Cost exploded and traffic did not

Requests flat, bill up 4×. Cost per request is not a constant — it is an emergent property of behaviour you did not change directly. The suspects, in order:

CauseMechanismCheck first
Cache silently stopped hittingCaching follows a strict hierarchy — tools → system → messages — and modifying earlier content invalidates that level and every level after it. Add one tool definition and your whole system-prompt cache is gone. Reads are 0.1× base input price; writes are 1.25× (5-min TTL) or 2× (1-hour). A cache that stops hitting does not cost 10× more than hitting — it costs 12.5× more, because you now write every time.cache_read_input_tokens vs cache_creation_input_tokens in the usage object. If reads collapsed to zero, find what you put in front of the prompt. Also check the 5-minute default TTL against your real traffic — low-traffic hours can miss every time.
Reasoning tokensReasoning tokens "are not visible via the API, they still occupy space in the model's context window and are billed as output tokens." A 500-token answer can carry thousands of invisible billed tokens. OpenAI suggests reserving at least 25,000 tokens for reasoning and output when starting out — that is the scale of what is hidden.output_tokens_details in the usage object. Compare visible answer length to billed output. If someone raised reasoning effort, that is a direct cost multiplier with no diff in your app.
Agent loop growthOne user request became 12 model calls instead of 3. Nothing in your traffic graph moved.Median and p99 calls per request, tracked over time. This metric almost never exists and is almost always the answer.
Conversation historyEvery turn resends the whole thread. Cost per conversation grows quadratically with turns. Users getting more engaged is the cost regression.Input tokens per call plotted against turn index. If it climbs linearly, you have no summarisation strategy.
Retrieval bloatSomeone raised K from 5 to 20 to fix a recall complaint. Input tokens quadrupled — and per lost-in-the-middle, quality may have dropped.K, and chunk size, and when they last changed. This is the fix that pays twice: cheaper and better.
Retry stormsA guardrail or parser rejects output and re-runs. A 20% failure rate is a 20% surcharge, invisible unless you count retries as calls.Validation failure rate. Retries billed but not logged as requests is the classic blind spot.

Caching mechanics and multipliers from Anthropic's prompt-caching documentation; reasoning-token billing and the 25,000-token reservation guidance from OpenAI's reasoning guide. Both were accurate when fetched — re-check before you budget on them. The Costing page covers the token math these failures distort.

Instrument

You cannot debug a cost regression you cannot decompose. Log tokens per call split into cached / uncached / reasoning / output, and calls per user request. Without those two, every cost investigation is archaeology.

The playbookThe first fifteen minutes

When someone says "the AI is broken" — in order, no skipping.

1) Get the exact input. Not a paraphrase. "It gets dates wrong" is not reproducible; a timestamp and a trace ID are. Most reports die here and should not.

2) Re-run it. If it does not reproduce, you are in section 4 — non-determinism, drift, or cache. Do not proceed as if it reproduces.

3) Read the retrieved context before the answer. The fork that decides everything: was the information there or not? Thirty seconds; saves a week.

4) Localise to a stage. Retrieval, assembly, generation, post-processing, or spec. Name it out loud before you change anything.

5) Check what changed outside your repo. Model alias, index rebuild, prompt edit, vendor deprecation, upstream data. Your diff is not the boundary of your system.

6) Fix the stage you localised — then add the case to the eval set. A bug you fixed without an eval case is a bug you will ship again. This step is the entire difference between a system that improves and one that oscillates.

What this page is really saying

  • Fluency is not correctnessand your monitoring cannot tell the difference. Only traces and humans can.
  • Most "model problems" are system problemschunking, ranking, verification, specification. MAST found failures cluster in design and coordination. Upgrading the model is the most expensive way to not fix them.
  • Every production failure is an eval case you did not have.That is the loop. There is no other loop.

ReferencesWhere these claims come from

Every external claim on this page is sourced. Where a number appears it came from the paper or the vendor doc, not from folklore.

Papers describe the failure; vendor docs describe the billing and the lifecycle. Both move. Re-check pricing, deprecation and caching pages quarterly — the mechanisms are stable, the numbers are not.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A cheaper Opus 5.5 reportedly changes agent-facing behavior, so swapping the model string can pass smoke tests and still fail in production.

    Add Claude Opus 5.5 to the Opus line and note that the cheaper model changed agent-facing behavior in ways smoke tests miss.

    Anthropic · 22 Sep 2026 · source

  • Updated this page Quantizing the KV cache can flip which experts a MoE model routes tokens to, because routing is discontinuous.

    On sub/pipe-mixture-experts-routing, note that MoE routing is discontinuous and that quantizing the KV cache can flip which experts fire.

    arXiv cs.AI · 12 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning