Mechanistic Interpretability
Mechanistic interpretability (mech interp) is a branch of AI safety and alignment research that tries to reverse-engineer the "circuits" inside transformers — treating the model as a decomposable compute graph and finding which attention head / MLP neuron does which specific task. For example:
- "Which head detects the start of a sentence?"
- "Which MLPs compute IOI (indirect object identification)?"
- "Is there a 'refuse-to-respond' circuit in the model?"
The goal is to turn the "black box" into a "white box" — knowing why the model answers as it does, not just what it answers.
Difference from "interpretability"
- Traditional interpretability: SHAP / LIME / feature importance — scores each input's contribution to the output
- Mechanistic interpretability: opens the model interior — locates the algorithmic-level compute units
The latter is more ambitious: if successful, you truly understand what the model is "thinking," not just which inputs influence the output.
Representative work
- Anthropic Transformer Circuits Thread (2021-): Chris Olah team's long-form series; the foundation of mech interp
- Induction Heads (Olsen et al. 2022): small models have a two-stage circuit that "copies earlier tokens"
- IOI Circuit (Wang et al. 2022): precisely located 15 attention heads in GPT-2 small that collectively perform "indirect object identification"
- Superposition / Polysemanticity: individual neurons encode multiple concepts
- Dictionary Learning (SAE / Sparse Autoencoders): decompose activations into large-scale sparse features — the mainstream 2024-2025 method
- Attribution Patching / Causal Tracing: causally localize circuits
Mainstream techniques
Probing
Train a simple classifier on hidden states to check whether a concept is encoded
Activation Patching
Swap "clean input" activations into "corrupted input" runs to identify which activation changes are causal
Circuit Discovery
Use attention maps + path search to find the minimal subgraph implementing a task
Sparse Autoencoders (SAE)
- Introduced by Anthropic in 2023
- Decompose high-dimensional activations into very large (million-scale) sparse features
- Each feature corresponds to a human-understandable concept
Why it matters
- AI safety: knowing when the model is "lying" / "hiding intent"
- Alignment research: internal mechanism of deception / power-seeking
- Model debugging: understanding failures / biases
- Trustworthy AI: regulators / medicine / law require "explainable" decisions
Limitations
- Only on small models so far: circuit work concentrates on GPT-2 / Claude Haiku (hundreds of millions of params); frontier models (GPT-4 / Claude Opus) are too large and complex
- Circuits aren't stable: fine-tuning / RLHF rearranges them
- No unified theory yet: many "local circuits," no whole-model-as- circuit-set theory
- Computationally heavy: a full circuit discovery takes thousands of GPU hours