Observability turns an opaque chain of LLM calls into a traceable, debuggable, measurable system.
ConceptWhat it is
Observability and tracing for LLM applications captures every step of a request, prompts, retrieved context, tool calls, intermediate reasoning, and final output, as a structured trace. It exists because LLM pipelines, especially multi-step agents and RAG systems, fail in ways that a simple input-output log cannot explain; understanding why an answer went wrong requires seeing every intermediate step.
Traces feed both live debugging and offline evaluation, turning production traffic into a source of continuous quality signal.
How it worksThe mechanics
Each call in a pipeline is wrapped with instrumentation that logs its inputs, outputs, latency, and token usage to a tracing backend, which links related calls into a single trace tree per user request, letting engineers replay any request step by step and attach evaluation scores or human labels after the fact.
At a glanceSee it
With a bad answer in hand, the trace lets you walk each span in order and localize the fault to retrieval, tools, prompt assembly, or the model itself — the diagnosis a plain input—output log cannot give.
Beyond the single request, traces roll up into a fleet-level loop — aggregate metrics and eval trends surface anomalies, drive a fix, and the next traces confirm it, with sampling as the escape valve on storage cost.
When to use itWhere it fits
- Multi-step agents or RAG pipelines where failures could originate at any stage.
- Production systems needing quality regression detection across model or prompt changes.
- Debugging user-reported issues that require seeing the exact prompt and context sent.
- Teams running continuous evaluation against a growing test set.
When NOT to use itLimits & anti-patterns
- Single-call, stateless prompts with no multi-step pipeline to trace.
- Early prototypes where instrumentation overhead outweighs the debugging benefit at that stage.
Trade-offsAdvantages & costs
Advantages
- Makes multi-step failures diagnosable instead of a black box.
- Enables systematic offline evaluation on real production traces.
- Surfaces cost and latency hotspots at the individual span level.
Trade-offs & costs
- Adds instrumentation code across every step of a pipeline.
- Trace storage and querying at scale has its own infrastructure cost.
- Logging full prompts and context raises the same data-sensitivity concerns as PII handling.
ExampleIn the real world
Instrumenting one of this site’s own kits, end to end. The kits call an OpenAI-compatible endpoint, so a single instrumentor covers every model call in the pipeline with no change to any call site:
# pip install arize-otel openinference-instrumentation-openai
from arize.otel import register
tracer_provider = register(
space_id="<your-space-id>", # Space Settings in the Arize app
api_key="<your-api-key>",
project_name="docs-qa", # one project per kit
)
from openinference.instrumentation.openai import OpenAIInstrumentor
OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)That is the whole hookup. Retrieval and parsing need one manual span each to appear as their own bars — the instrumentor sees model calls, not your functions:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("retrieve") as span:
hits = index.search(question, k=5)
span.set_attribute("openinference.span.kind", "RETRIEVER")
span.set_attribute("retrieval.k", 5)Why it is worth the two lines. This site’s own kits publish latency as a median and a p95 read from their run records — enough to show that one kit spends 99.6% of its call inside the model and 18 ms in retrieval, and that is the end of what an aggregate can say. A trace answers the next question: which step, on which request, and what did it cost. A percentile cannot tell you whether a slow p95 is one bad call or a fat tail; a span tree shows you.
Swap the destination, not the code. The spans above are OpenInference, which is an OpenTelemetry convention rather than a vendor format — the same instrumentation sends to a hosted backend, to a Phoenix instance running on your own machine with no account, or to a file you commit. Instrument first, choose the backend second; that ordering is what keeps the choice reversible.
ToolsHow to implement it
- OpenTelemetry + OpenInferencethe vendor-neutral span format every tool below reads. Instrumenting to this standard is what makes the backend a reversible decision rather than a coupling.
- Arize AXhosted tracing with per-span latency, tokens and cost, plus session and agent-path views.
- Arize Phoenixthe open-source sibling; same span format, runs locally, no account, so a fork of a project can be traced by whoever cloned it.
- Langfuseopen-source tracing and evaluation, self-hostable.
- LangSmithtracing and evaluation built for LangChain and LangGraph pipelines.
- Weights & Biases Weavetracing and evaluation integrated with experiment tracking.
- Datadog LLM Observabilityenterprise APM extended to LLM-specific spans and metrics.
Cost & effortWhat it takes
Tracing overhead is usually low single-digit percent latency and minor storage cost; the bigger investment is engineering time instrumenting every pipeline step consistently.