Home › Neural Networks & the Transformer › Encoder-Decoder
🧠 · Foundations

Encoder-Decoder

The original transformer design: read the input fully, then generate an output conditioned on it.

In one line

Encoder-decoder models read the full source with one stack and generate the output with another, ideal for translation and summarization.

ConceptWhat it is

An encoder-decoder transformer uses two stacks: an encoder that reads the entire input bidirectionally to build a rich representation, and a decoder that generates the output autoregressively while cross-attending to that encoded representation. It exists for sequence-to-sequence tasks where the output is a transformation of a distinct input, like translating a sentence or summarizing a document, rather than open-ended continuation.

This was the original transformer design from the 2017 paper, and it remains the natural fit whenever there is a clear source sequence and target sequence that are meaningfully different.

How it worksThe mechanics

The encoder processes the full source sequence with bidirectional self-attention, producing a set of context vectors; the decoder then generates the target sequence token by token, using causal self-attention over what it has generated so far plus cross-attention over the encoder's output, so each generated token can pull relevant information from anywhere in the source.

At a glanceSee it

Encoder-Decoder diagram
Encoder-Decoder diagram 1

Zooms inside the single cross-attention box to reveal that an encoder-decoder actually runs three distinct attention types — and only the decoder’s self-attention is causally masked.

Encoder-Decoder diagram 2

Places encoder-decoder within the wider architecture choice, showing when a two-stack design beats an encoder-only or decoder-only model.

When to use itWhere it fits

  • Machine translation between two languages.
  • Abstractive summarization of a fixed source document.
  • Structured transformation tasks like text-to-SQL or data-to-text generation.
  • Speech-to-text or other tasks with a clearly distinct input and output modality or format.

When NOT to use itLimits & anti-patterns

  • Open-ended chat or free-form generation without a fixed source to condition on, where a decoder-only model is simpler and equally effective.
  • Pure understanding tasks like classification, where an encoder-only model is cheaper.
  • Very large-scale general-purpose assistants, where the industry has largely converged on decoder-only architectures for flexibility.

Trade-offsAdvantages & costs

Advantages
  • Naturally suited to input-to-output transformation tasks with a clear source and target.
  • Encoder's bidirectional context improves grounding in the source document.
  • Cross-attention gives the decoder direct access to any part of the source.
  • Well-established, battle-tested architecture for translation and summarization.
Trade-offs & costs
  • More complex and parameter-heavy than a single-stack decoder-only model.
  • Less flexible for general-purpose, open-ended assistant use cases.
  • Largely displaced in mindshare by decoder-only LLMs that handle seq2seq tasks via prompting.
  • Training requires paired input-output data, which can be harder to source at scale.

ExampleIn the real world

Google Translate's neural backend and the T5 and BART model families use encoder-decoder transformers for translation and summarization tasks.

ToolsHow to implement it

  • Hugging Face Transformershosts T5, BART, and mT5 encoder-decoder checkpoints.
  • MarianMT / OPUS-MTopen translation models built on this architecture.
  • FairseqMeta's sequence-to-sequence modeling toolkit.
  • Google Cloud Translation APIproduction encoder-decoder translation service.

Cost & effortWhat it takes

Training cost is moderate to high depending on scale; inference requires running both stacks, adding some latency versus decoder-only; needs paired source-target data for fine-tuning.

A living map of modern AI — kept current every morning