Self-Attention
Self-Attention is the core computational unit of the Transformer architecture. It lets every position in a sequence "see" every other position and assign different weights by relevance.
How it works
Given three matrices derived from the input:
- Query
- Key
- Value
The attention formula:
Attention(Q, K, V) = softmax(QK^T / √d_k) V
Intuition: each position uses Q to "ask" every position's K, computes relevance scores (dividing by √d_k prevents softmax saturation), then weighted-sums V.
Complexity
Self-attention time and space are both O(n²) in sequence length. This means:
- Long documents / long contexts quickly eat up GPU memory.
- It's the pain point of long-context research — see Mamba and other sub-quadratic architectures.
Multi-Head Attention
In practice, models don't do attention once — they split Q/K/V into multiple "heads", each computes its own attention, then concatenate. This lets the model attend to different sub-spaces of relationships simultaneously.