🤖 · Models

Chat

The multi-turn, conversational tuning layer that makes a model feel like an assistant.

In one line

Chat models are tuned specifically for back-and-forth conversation, remembering context across many turns.

ConceptWhat it is

A chat model is an instruction-tuned model further specialized for multi-turn dialogue, trained on conversational data so it maintains context across turns, adopts a consistent persona, and handles follow-up questions naturally. It exists because single-turn instruction following is not enough for the conversational, iterative way people actually want to interact with an assistant.

Chat tuning typically adds a structured format, system, user, and assistant roles, and trains the model to behave well within that structure across long, evolving conversations.

How it worksThe mechanics

The model is fine-tuned on multi-turn conversation transcripts formatted with explicit role markers, learning to track what was said earlier in the conversation, resolve references like pronouns to prior turns, and stay consistent with any system-level instructions; reinforcement learning from human feedback further shapes tone, helpfulness, and refusal behavior across turns.

At a glanceSee it

Chat diagram
Chat diagram 1

Under the hood a chat template flattens the role-tagged message list into one delimited token stream, which the model reads and continues until the end-of-turn marker — and a wrong template garbles the roles.

Chat diagram 2

When accumulated history outgrows the context window the app must decide how to shrink it — drop, summarize, or pin — each trading tokens against how much the model still remembers.

When to use itWhere it fits

  • Customer support and conversational assistant products.
  • Multi-turn agents that need to maintain context across a session.
  • Interactive coding assistants where follow-up refinements build on prior turns.
  • Any consumer-facing product where a conversational interface is the primary UX.

When NOT to use itLimits & anti-patterns

  • Single-shot batch tasks like bulk classification, where the conversational overhead adds no value.
  • Latency-critical, stateless API calls where carrying conversation history adds unnecessary token cost.
  • Highly structured extraction pipelines better served by a plain instruction-tuned model without persona shaping.

Trade-offsAdvantages & costs

Advantages
  • Naturally handles multi-turn context and follow-up questions.
  • Consistent persona and tone across a conversation improves user trust.
  • Widely supported by chat-completion APIs and tooling.
  • Well suited to the assistant and agent product patterns dominating 2026 AI products.
Trade-offs & costs
  • Longer conversation history increases token cost and latency over a session.
  • Persona and safety tuning can make responses feel more guarded or verbose.
  • Context window limits can cause the model to lose track of earlier turns.
  • Not optimized for pure single-shot precision tasks the way a narrow instruction model might be.

ExampleIn the real world

ChatGPT, Claude.ai, and Google's Gemini app are all chat-tuned models wrapped in a conversational interface that maintains session context across many turns.

ToolsHow to implement it

  • OpenAI Chat Completions / Responses APIstandard interface for multi-turn chat models.
  • Anthropic Messages APIrole-structured conversation format for Claude.
  • LangChain memory modulesmanages conversation history and summarization.
  • Redis / vector storespersist long-running chat session state.

Cost & effortWhat it takes

Cost scales with cumulative conversation tokens, not just the latest turn; latency grows as history lengthens unless truncated or summarized; needs conversational training data to build, but is available off the shelf from every major provider.

A living map of modern AI — kept current every morning