🧠 · Foundations

Decoder (GPT)

The generative, left-to-right transformer behind ChatGPT and nearly every modern chat model.

In one line

Decoder-only models generate text one token at a time, each token seeing only what came before it.

ConceptWhat it is

A decoder-only transformer such as GPT generates text autoregressively, predicting each next token based only on the tokens that came before it, using a causal attention mask that blocks any peek at future tokens. It exists to model open-ended generation: writing, chatting, reasoning, and coding, where the model must produce new content rather than just understand existing text.

This simplicity, one objective, next-token prediction, at massive scale, turned out to be the recipe behind essentially every modern chat assistant and coding model.

How it worksThe mechanics

The model processes a prompt token by token, and at each generation step it attends only to previous tokens, computing a probability distribution over the vocabulary for the next token, which is sampled or chosen greedily and appended before the process repeats. Training uses the same mechanism at scale: predict the next token across trillions of tokens of text, which implicitly teaches grammar, facts, and reasoning patterns.

At a glanceSee it

Decoder (GPT) diagram
Decoder (GPT) diagram 1

Inside the single sample-or-pick step — temperature reshapes the distribution, then a greedy-or-sampled branch with top-k or top-p truncation settles on the next token.

Decoder (GPT) diagram 2

Why sequential decoding still runs fast — a one-time prefill fills the key-value cache, then each decode step computes only the newest token and appends it, rather than recomputing the whole prompt.

When to use itWhere it fits

  • Chatbots, assistants, and any open-ended text or code generation task.
  • Few-shot and zero-shot prompting without task-specific fine-tuning.
  • Reasoning-style tasks where the model needs to think through steps before answering.
  • Building agents that plan, call tools, and produce free-form outputs.

When NOT to use itLimits & anti-patterns

  • Pure classification or embedding tasks, where a smaller encoder model is cheaper and often more accurate.
  • Tasks demanding strict bidirectional understanding of a fixed input, like span extraction.
  • Extremely latency-sensitive, high-throughput scoring tasks where autoregressive generation is overkill.

Trade-offsAdvantages & costs

Advantages
  • Extremely flexible: one model handles chat, coding, summarization, and reasoning.
  • Strong zero-shot and few-shot performance from large-scale pretraining.
  • Simple training objective that scales predictably with data and compute.
  • Massive ecosystem of tooling, fine-tuning methods, and deployment options.
Trade-offs & costs
  • Autoregressive generation is inherently sequential and slower per token than encoder inference.
  • Can hallucinate fluent but incorrect content.
  • Compute and memory cost grow with both model size and context length.
  • Harder to control precisely than a fine-tuned narrow classifier.

ExampleIn the real world

OpenAI's GPT-4 and GPT-5 family, Anthropic's Claude, and Meta's Llama are all decoder-only transformers powering chat products used by hundreds of millions of people.

ToolsHow to implement it

  • vLLMhigh-throughput serving engine for decoder-only LLMs.
  • Hugging Face Transformersloads and fine-tunes GPT-style checkpoints.
  • OpenAI/Anthropic APIsproduction access to frontier decoder models.
  • llama.cppruns open-weight decoder models efficiently on local hardware.

Cost & effortWhat it takes

Frontier models cost tens of millions to pretrain; inference cost scales with tokens generated and context length; latency is dominated by sequential decoding unless batched or speculative decoding is used.

A living map of modern AI — kept current every morning