Home › Agents & Tool Use › API response
🕹️ · Build

API response

Returning a model's answer as a structured, synchronous response another system can rely on.

In one line

The moment you return JSON instead of prose, the model's output becomes a contract, and contracts need validation rather than hope.

ConceptWhat it is

An API response is the plainest exit a model's answer can take: a caller asks, waits, and receives a structured result. It is the shape most systems integrate against, and the one where the gap between a language model and a dependable component is most visible.

The difficulty is that a schema is a promise about every response, and a model is a probabilistic process. Most of the engineering here is closing that gap — constraining output, validating it, and deciding what to return when the model produces something the schema cannot accept.

How it worksThe mechanics

The endpoint declares a schema. The model is asked for output conforming to it, using structured-output or tool-call modes where the provider offers them, since those constrain decoding rather than merely requesting good behaviour. The result is parsed and validated before it leaves the process.

Validation failure is a designed path, not an exception. A single constrained retry with the validation error fed back resolves most cases; beyond that the endpoint returns an explicit error rather than a plausible object, because a caller that receives malformed data it cannot detect is worse off than one that receives a failure.

At a glanceSee it

API response diagram

Validation is a designed path with two exits. Returning an explicit failure is better than returning a plausible object the caller cannot check.

When to use itWhere it fits

  • When another system consumes the output and needs fields rather than paragraphs.
  • Extraction, classification and routing, where the answer is a value and not an explanation.
  • Anywhere the response feeds a database write, a decision rule or another service.
  • When you need latency bounded and a timeout that means something to the caller.

When NOT to use itLimits & anti-patterns

  • Long generative answers a person will read, where waiting for the whole response feels broken and streaming is the right exit.
  • Work that legitimately takes minutes, which belongs in a job with a callback rather than an open connection.
  • When the caller can accept partial or best-effort output; a strict schema then costs more than it returns.
  • As a way to make an unreliable model reliable — the schema constrains shape, never correctness.

Trade-offsAdvantages & costs

Advantages
  • Integrates with everything, because a typed HTTP response is the universal interface.
  • Validation makes malformed output a caught error rather than a downstream mystery.
  • Easy to test, cache and rate-limit with ordinary infrastructure.
  • Latency and cost per call are simple to attribute, which makes the economics legible.
Trade-offs & costs
  • The caller waits for the whole generation, so perceived latency is the full response time.
  • Schema conformance is not correctness, and a well-formed wrong answer passes every check.
  • Retries multiply cost and latency exactly when the system is already struggling.
  • Strict schemas can suppress the model's ability to signal uncertainty unless a field exists for it.

ExampleIn the real world

An invoice endpoint accepts a PDF and returns supplier, date, currency and total. The model reads a scanned document and returns a total with the thousands separator interpreted as a decimal point. The schema is satisfied and the number is wrong by three orders of magnitude, which is why the endpoint also carries a confidence field and a range check — the parts that catch what validation cannot.

ToolsHow to implement it

  • Provider structured-output modesconstrained decoding against a JSON schema, which is stronger than asking politely in the prompt.
  • Pydantic or Zodone schema definition that both constrains the model and validates the result.
  • FastAPI or Expressordinary web framework, since nothing about this needs special infrastructure.
  • An idempotency keyso a client retry after a timeout does not double the work or the charge.

Cost & effortWhat it takes

One model call plus occasional retries; cost is dominated by output tokens and the retry rate, so measure the rate before assuming it is small. Latency is the full generation, typically one to several seconds. Engineering effort is low, and the part that takes real time is the validation and failure design rather than the endpoint.

A living map of modern AI — kept current every morning