🧠 · Foundations

Transformer

The 2017 architecture that replaced recurrence with attention and unlocked modern LLMs.

In one line

The transformer lets every token look directly at every other token in parallel, which is why it scales so well.

ConceptWhat it is

The Transformer, introduced in the 2017 paper "Attention Is All You Need," replaced step-by-step recurrence with self-attention, letting every token in a sequence weigh and combine information from every other token in a single parallel operation. It exists because recurrent networks were slow to train and struggled with long-range dependencies; attention solves both by computing relationships directly, regardless of distance.

This design is the common ancestor of essentially every modern large language and multimodal model, from GPT to BERT to vision transformers, because it scales predictably with more data and compute.

How it worksThe mechanics

Input tokens are embedded into vectors and each token produces a query, key, and value; attention scores are computed between every pair of tokens by comparing queries to keys, and the resulting weights are used to blend value vectors into a new, context-aware representation for each token, all in parallel across the sequence. Stacking many of these attention-plus-feedforward blocks, with residual connections and normalization, produces deep representations that are then used for generation, classification, or embedding.

At a glanceSee it

Transformer diagram
Transformer diagram 1

One transformer block wraps attention in residual connections, normalization, and a feed-forward layer, and the same block is stacked many times to build depth.

Transformer diagram 2

Whether a task calls for encoder-only, decoder-only, or encoder-decoder depends on if you need to understand, generate, or transform a sequence — with decoders masking future tokens so they cannot peek at the answer.

When to use itWhere it fits

  • Any large-scale language, code, or multimodal modeling task, given its dominance across NLP and beyond.
  • Tasks requiring long-range context, since attention connects distant tokens directly.
  • Building on top of pretrained transformer checkpoints via fine-tuning or adapters.
  • Scenarios where parallel, GPU-friendly training matters for iteration speed.

When NOT to use itLimits & anti-patterns

  • Extremely long sequences at naive quadratic attention cost, where sparse or linear-attention variants are needed instead.
  • Tiny, resource-constrained deployments where a small CNN or classical model is cheaper and sufficient.
  • Situations demanding strict recurrence semantics, like certain streaming control systems with hard real-time causal constraints.

Trade-offsAdvantages & costs

Advantages
  • Parallelizable training that scales efficiently with more GPUs and data.
  • Captures long-range dependencies directly through attention.
  • One architecture generalizes across text, vision, audio, and multimodal tasks.
  • Massive ecosystem of pretrained checkpoints and tooling.
Trade-offs & costs
  • Self-attention cost grows quadratically with sequence length.
  • Requires large amounts of data and compute to train from scratch.
  • Memory-hungry at long context lengths without optimizations like FlashAttention.
  • Opaque internals make behavior harder to interpret than simpler models.

ExampleIn the real world

OpenAI's GPT series, Google's BERT and Gemini, and Meta's Llama models are all built on the transformer architecture, as is Stable Diffusion's text encoder.

ToolsHow to implement it

  • Hugging Face Transformersthe standard library for pretrained transformer models.
  • PyTorchthe dominant framework for training transformer architectures.
  • FlashAttentionmemory-efficient attention kernels for long context.
  • JAX/Flaxused by Google DeepMind for large-scale transformer training.

Cost & effortWhat it takes

Pretraining from scratch costs millions of dollars in compute; fine-tuning or using pretrained checkpoints is comparatively cheap; inference cost and latency scale with context length and model size.

A living map of modern AI — kept current every morning