Home › Deployment, Inference & LLMOps › Observability / tracing
⚙️ · Operate

Observability / tracing

Instrumenting LLM calls end-to-end to debug, evaluate, and monitor production behavior.

In one line

Observability turns an opaque chain of LLM calls into a traceable, debuggable, measurable system.

ConceptWhat it is

Observability and tracing for LLM applications captures every step of a request, prompts, retrieved context, tool calls, intermediate reasoning, and final output, as a structured trace. It exists because LLM pipelines, especially multi-step agents and RAG systems, fail in ways that a simple input-output log cannot explain; understanding why an answer went wrong requires seeing every intermediate step.

Traces feed both live debugging and offline evaluation, turning production traffic into a source of continuous quality signal.

How it worksThe mechanics

Each call in a pipeline is wrapped with instrumentation that logs its inputs, outputs, latency, and token usage to a tracing backend, which links related calls into a single trace tree per user request, letting engineers replay any request step by step and attach evaluation scores or human labels after the fact.

At a glanceSee it

Observability / tracing diagram
Observability / tracing diagram 1

With a bad answer in hand, the trace lets you walk each span in order and localize the fault to retrieval, tools, prompt assembly, or the model itself — the diagnosis a plain input—output log cannot give.

Observability / tracing diagram 2

Beyond the single request, traces roll up into a fleet-level loop — aggregate metrics and eval trends surface anomalies, drive a fix, and the next traces confirm it, with sampling as the escape valve on storage cost.

When to use itWhere it fits

  • Multi-step agents or RAG pipelines where failures could originate at any stage.
  • Production systems needing quality regression detection across model or prompt changes.
  • Debugging user-reported issues that require seeing the exact prompt and context sent.
  • Teams running continuous evaluation against a growing test set.

When NOT to use itLimits & anti-patterns

  • Single-call, stateless prompts with no multi-step pipeline to trace.
  • Early prototypes where instrumentation overhead outweighs the debugging benefit at that stage.

Trade-offsAdvantages & costs

Advantages
  • Makes multi-step failures diagnosable instead of a black box.
  • Enables systematic offline evaluation on real production traces.
  • Surfaces cost and latency hotspots at the individual span level.
Trade-offs & costs
  • Adds instrumentation code across every step of a pipeline.
  • Trace storage and querying at scale has its own infrastructure cost.
  • Logging full prompts and context raises the same data-sensitivity concerns as PII handling.

ExampleIn the real world

Instrumenting one of this site’s own kits, end to end. The kits call an OpenAI-compatible endpoint, so a single instrumentor covers every model call in the pipeline with no change to any call site:

code
# pip install arize-otel openinference-instrumentation-openai

from arize.otel import register

tracer_provider = register(
    space_id="<your-space-id>",     # Space Settings in the Arize app
    api_key="<your-api-key>",
    project_name="docs-qa",          # one project per kit
)

from openinference.instrumentation.openai import OpenAIInstrumentor

OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)

That is the whole hookup. Retrieval and parsing need one manual span each to appear as their own bars — the instrumentor sees model calls, not your functions:

code
from opentelemetry import trace

tracer = trace.get_tracer(__name__)

with tracer.start_as_current_span("retrieve") as span:
    hits = index.search(question, k=5)
    span.set_attribute("openinference.span.kind", "RETRIEVER")
    span.set_attribute("retrieval.k", 5)

Why it is worth the two lines. This site’s own kits publish latency as a median and a p95 read from their run records — enough to show that one kit spends 99.6% of its call inside the model and 18 ms in retrieval, and that is the end of what an aggregate can say. A trace answers the next question: which step, on which request, and what did it cost. A percentile cannot tell you whether a slow p95 is one bad call or a fat tail; a span tree shows you.

Swap the destination, not the code. The spans above are OpenInference, which is an OpenTelemetry convention rather than a vendor format — the same instrumentation sends to a hosted backend, to a Phoenix instance running on your own machine with no account, or to a file you commit. Instrument first, choose the backend second; that ordering is what keeps the choice reversible.

ToolsHow to implement it

  • OpenTelemetry + OpenInferencethe vendor-neutral span format every tool below reads. Instrumenting to this standard is what makes the backend a reversible decision rather than a coupling.
  • Arize AXhosted tracing with per-span latency, tokens and cost, plus session and agent-path views.
  • Arize Phoenixthe open-source sibling; same span format, runs locally, no account, so a fork of a project can be traced by whoever cloned it.
  • Langfuseopen-source tracing and evaluation, self-hostable.
  • LangSmithtracing and evaluation built for LangChain and LangGraph pipelines.
  • Weights & Biases Weavetracing and evaluation integrated with experiment tracking.
  • Datadog LLM Observabilityenterprise APM extended to LLM-specific spans and metrics.

Cost & effortWhat it takes

Tracing overhead is usually low single-digit percent latency and minor storage cost; the bigger investment is engineering time instrumenting every pipeline step consistently.

A living map of modern AI — kept current every morning