Skip to main content
ZICQ

Wiki AI Concepts

Mechanistic Interpretability

AI Concepts
Aliases: Mechanistic Interpretability Interpretability Model Interpretability Circuits ·2026-09-19

Mechanistic Interpretability

Mechanistic interpretability (mech interp) is a branch of AI safety and alignment research that tries to reverse-engineer the "circuits" inside transformers — treating the model as a decomposable compute graph and finding which attention head / MLP neuron does which specific task. For example:

  • "Which head detects the start of a sentence?"
  • "Which MLPs compute IOI (indirect object identification)?"
  • "Is there a 'refuse-to-respond' circuit in the model?"

The goal is to turn the "black box" into a "white box" — knowing why the model answers as it does, not just what it answers.

Difference from "interpretability"

  • Traditional interpretability: SHAP / LIME / feature importance — scores each input's contribution to the output
  • Mechanistic interpretability: opens the model interior — locates the algorithmic-level compute units

The latter is more ambitious: if successful, you truly understand what the model is "thinking," not just which inputs influence the output.

Representative work

  • Anthropic Transformer Circuits Thread (2021-): Chris Olah team's long-form series; the foundation of mech interp
  • Induction Heads (Olsen et al. 2022): small models have a two-stage circuit that "copies earlier tokens"
  • IOI Circuit (Wang et al. 2022): precisely located 15 attention heads in GPT-2 small that collectively perform "indirect object identification"
  • Superposition / Polysemanticity: individual neurons encode multiple concepts
  • Dictionary Learning (SAE / Sparse Autoencoders): decompose activations into large-scale sparse features — the mainstream 2024-2025 method
  • Attribution Patching / Causal Tracing: causally localize circuits

Mainstream techniques

Probing

Train a simple classifier on hidden states to check whether a concept is encoded

Activation Patching

Swap "clean input" activations into "corrupted input" runs to identify which activation changes are causal

Circuit Discovery

Use attention maps + path search to find the minimal subgraph implementing a task

Sparse Autoencoders (SAE)

  • Introduced by Anthropic in 2023
  • Decompose high-dimensional activations into very large (million-scale) sparse features
  • Each feature corresponds to a human-understandable concept

Why it matters

  • AI safety: knowing when the model is "lying" / "hiding intent"
  • Alignment research: internal mechanism of deception / power-seeking
  • Model debugging: understanding failures / biases
  • Trustworthy AI: regulators / medicine / law require "explainable" decisions

Limitations

  • Only on small models so far: circuit work concentrates on GPT-2 / Claude Haiku (hundreds of millions of params); frontier models (GPT-4 / Claude Opus) are too large and complex
  • Circuits aren't stable: fine-tuning / RLHF rearranges them
  • No unified theory yet: many "local circuits," no whole-model-as- circuit-set theory
  • Computationally heavy: a full circuit discovery takes thousands of GPU hours