Home › Neural Networks & the Transformer › Key term › Self-attention
Key term · Foundations

Self-attention

The mechanism letting every token weigh every other token when building meaning.

In one line

Self-attention lets each word look at every other word to figure out what it means in context.

DefinitionWhat it means

Self-attention is the core mechanism of the transformer architecture: for every token in a sequence, the model computes how much attention to pay to every other token, producing a weighted, context-aware representation rather than a fixed, isolated one. This is what lets a transformer resolve that bank means a riverbank in one sentence and a financial institution in another, by attending differently to surrounding words each time.

Why it mattersWhy you should care

Self-attention is the innovation, introduced in the 2017 Attention Is All You Need paper, that replaced recurrent networks and made large-scale parallel training of language models practical, directly enabling the current generation of LLMs. Its computational cost, which grows quadratically with sequence length, is also the reason context-window limits and efficient-attention research remain active engineering concerns for anyone deploying long-context applications.

At a glanceSee it

Self-attention diagram
Self-attention diagram 1

Opens the attention-weight black box — scores are scaled dot products between a query and every key, softmaxed into weights, then used to average the values.

Self-attention diagram 2

Multi-head attention runs several attention computations in parallel so different heads specialize in different relationships before the results are merged.

Where you see itIn the wild

  • Transformer architecture diagrams in a model paper
  • Attention-map visualizations explaining model behavior
  • Discussions of quadratic cost limiting context length
A living map of modern AI — kept current every morning