Decoder-only models generate text one token at a time, each token seeing only what came before it.
ConceptWhat it is
A decoder-only transformer such as GPT generates text autoregressively, predicting each next token based only on the tokens that came before it, using a causal attention mask that blocks any peek at future tokens. It exists to model open-ended generation: writing, chatting, reasoning, and coding, where the model must produce new content rather than just understand existing text.
This simplicity, one objective, next-token prediction, at massive scale, turned out to be the recipe behind essentially every modern chat assistant and coding model.
How it worksThe mechanics
The model processes a prompt token by token, and at each generation step it attends only to previous tokens, computing a probability distribution over the vocabulary for the next token, which is sampled or chosen greedily and appended before the process repeats. Training uses the same mechanism at scale: predict the next token across trillions of tokens of text, which implicitly teaches grammar, facts, and reasoning patterns.
At a glanceSee it
Inside the single sample-or-pick step — temperature reshapes the distribution, then a greedy-or-sampled branch with top-k or top-p truncation settles on the next token.
Why sequential decoding still runs fast — a one-time prefill fills the key-value cache, then each decode step computes only the newest token and appends it, rather than recomputing the whole prompt.
When to use itWhere it fits
- Chatbots, assistants, and any open-ended text or code generation task.
- Few-shot and zero-shot prompting without task-specific fine-tuning.
- Reasoning-style tasks where the model needs to think through steps before answering.
- Building agents that plan, call tools, and produce free-form outputs.
When NOT to use itLimits & anti-patterns
- Pure classification or embedding tasks, where a smaller encoder model is cheaper and often more accurate.
- Tasks demanding strict bidirectional understanding of a fixed input, like span extraction.
- Extremely latency-sensitive, high-throughput scoring tasks where autoregressive generation is overkill.
Trade-offsAdvantages & costs
Advantages
- Extremely flexible: one model handles chat, coding, summarization, and reasoning.
- Strong zero-shot and few-shot performance from large-scale pretraining.
- Simple training objective that scales predictably with data and compute.
- Massive ecosystem of tooling, fine-tuning methods, and deployment options.
Trade-offs & costs
- Autoregressive generation is inherently sequential and slower per token than encoder inference.
- Can hallucinate fluent but incorrect content.
- Compute and memory cost grow with both model size and context length.
- Harder to control precisely than a fine-tuned narrow classifier.
ExampleIn the real world
OpenAI's GPT-4 and GPT-5 family, Anthropic's Claude, and Meta's Llama are all decoder-only transformers powering chat products used by hundreds of millions of people.ToolsHow to implement it
- vLLMhigh-throughput serving engine for decoder-only LLMs.
- Hugging Face Transformersloads and fine-tunes GPT-style checkpoints.
- OpenAI/Anthropic APIsproduction access to frontier decoder models.
- llama.cppruns open-weight decoder models efficiently on local hardware.
Cost & effortWhat it takes
Frontier models cost tens of millions to pretrain; inference cost scales with tokens generated and context length; latency is dominated by sequential decoding unless batched or speculative decoding is used.