Skip to main content
ZICQ

Wiki Concepts

Self-Attention

Concepts
Aliases: attention Self-Attention ·2026-09-14

Self-Attention

Self-Attention is the core computational unit of the Transformer architecture. It lets every position in a sequence "see" every other position and assign different weights by relevance.

How it works

Given three matrices derived from the input:

  • Query
  • Key
  • Value

The attention formula:

Attention(Q, K, V) = softmax(QK^T / √d_k) V

Intuition: each position uses Q to "ask" every position's K, computes relevance scores (dividing by √d_k prevents softmax saturation), then weighted-sums V.

Complexity

Self-attention time and space are both O(n²) in sequence length. This means:

  • Long documents / long contexts quickly eat up GPU memory.
  • It's the pain point of long-context research — see Mamba and other sub-quadratic architectures.

Multi-Head Attention

In practice, models don't do attention once — they split Q/K/V into multiple "heads", each computes its own attention, then concatenate. This lets the model attend to different sub-spaces of relationships simultaneously.