Self-attention lets each word look at every other word to figure out what it means in context.
DefinitionWhat it means
Self-attention is the core mechanism of the transformer architecture: for every token in a sequence, the model computes how much attention to pay to every other token, producing a weighted, context-aware representation rather than a fixed, isolated one. This is what lets a transformer resolve that bank means a riverbank in one sentence and a financial institution in another, by attending differently to surrounding words each time.
Why it mattersWhy you should care
Self-attention is the innovation, introduced in the 2017 Attention Is All You Need paper, that replaced recurrent networks and made large-scale parallel training of language models practical, directly enabling the current generation of LLMs. Its computational cost, which grows quadratically with sequence length, is also the reason context-window limits and efficient-attention research remain active engineering concerns for anyone deploying long-context applications.
At a glanceSee it
Opens the attention-weight black box — scores are scaled dot products between a query and every key, softmaxed into weights, then used to average the values.
Multi-head attention runs several attention computations in parallel so different heads specialize in different relationships before the results are merged.
Where you see itIn the wild
- Transformer architecture diagrams in a model paper
- Attention-map visualizations explaining model behavior
- Discussions of quadratic cost limiting context length