Encoder-decoder models read the full source with one stack and generate the output with another, ideal for translation and summarization.
ConceptWhat it is
An encoder-decoder transformer uses two stacks: an encoder that reads the entire input bidirectionally to build a rich representation, and a decoder that generates the output autoregressively while cross-attending to that encoded representation. It exists for sequence-to-sequence tasks where the output is a transformation of a distinct input, like translating a sentence or summarizing a document, rather than open-ended continuation.
This was the original transformer design from the 2017 paper, and it remains the natural fit whenever there is a clear source sequence and target sequence that are meaningfully different.
How it worksThe mechanics
The encoder processes the full source sequence with bidirectional self-attention, producing a set of context vectors; the decoder then generates the target sequence token by token, using causal self-attention over what it has generated so far plus cross-attention over the encoder's output, so each generated token can pull relevant information from anywhere in the source.
At a glanceSee it
Zooms inside the single cross-attention box to reveal that an encoder-decoder actually runs three distinct attention types — and only the decoder’s self-attention is causally masked.
Places encoder-decoder within the wider architecture choice, showing when a two-stack design beats an encoder-only or decoder-only model.
When to use itWhere it fits
- Machine translation between two languages.
- Abstractive summarization of a fixed source document.
- Structured transformation tasks like text-to-SQL or data-to-text generation.
- Speech-to-text or other tasks with a clearly distinct input and output modality or format.
When NOT to use itLimits & anti-patterns
- Open-ended chat or free-form generation without a fixed source to condition on, where a decoder-only model is simpler and equally effective.
- Pure understanding tasks like classification, where an encoder-only model is cheaper.
- Very large-scale general-purpose assistants, where the industry has largely converged on decoder-only architectures for flexibility.
Trade-offsAdvantages & costs
Advantages
- Naturally suited to input-to-output transformation tasks with a clear source and target.
- Encoder's bidirectional context improves grounding in the source document.
- Cross-attention gives the decoder direct access to any part of the source.
- Well-established, battle-tested architecture for translation and summarization.
Trade-offs & costs
- More complex and parameter-heavy than a single-stack decoder-only model.
- Less flexible for general-purpose, open-ended assistant use cases.
- Largely displaced in mindshare by decoder-only LLMs that handle seq2seq tasks via prompting.
- Training requires paired input-output data, which can be harder to source at scale.
ExampleIn the real world
Google Translate's neural backend and the T5 and BART model families use encoder-decoder transformers for translation and summarization tasks.ToolsHow to implement it
- Hugging Face Transformershosts T5, BART, and mT5 encoder-decoder checkpoints.
- MarianMT / OPUS-MTopen translation models built on this architecture.
- FairseqMeta's sequence-to-sequence modeling toolkit.
- Google Cloud Translation APIproduction encoder-decoder translation service.
Cost & effortWhat it takes
Training cost is moderate to high depending on scale; inference requires running both stacks, adding some latency versus decoder-only; needs paired source-target data for fine-tuning.