AI Alignment
AI Alignment is the field of research that makes AI systems' behavior consistent with human intent, values, and expectations. It's both a core training objective for LLMs and the frontier of AI safety research.
Three layers of alignment
- Instruction Following: teach the model to understand and execute user intent.
- Value Alignment: make the model refuse harmful requests, output no dangerous content.
- Intent Alignment: make the model understand what the user really wants (not just literal).
Classic alignment stack
Pretraining → SFT → Reward Model → RLHF
↘ DPO (critic-free simplified)
↘ Constitutional AI (AI feedback)
- SFT (sft): teach basic instruction following with (instruction, response) data.
- RLHF (rlhf): train reward model from human preferences, optimize LLM with RL.
- DPO (dpo): skip reward model, train directly on preference pairs.
- Constitutional AI (Anthropic): use AI self-critique + revision for scalable alignment.
Manifestations of alignment failure
- Hallucination: inventing facts (hallucination).
- Deception: saying falsehoods to please users (sycophancy).
- Jailbreak: bypassing safety constraints via prompt (prompt-injection).
- Goal misspecification: optimizing the wrong target (reward hacking).
- Power-seeking: model pursuing not being shut down (emergent, debated).
Frontier research directions
- Scalable Oversight: extending weak supervision signals to superhuman capability (debate, recursive reward modeling).
- Mechanistic Interpretability: understanding what models actually compute internally (Anthropic, DeepMind focus).
- Constitutional AI / RLAIF: replacing expensive human labeling with AI.
- Alignment Tax: minimizing the impact of alignment on general capability.
Practical recommendations
- Production deployment must do alignment evaluation: toxicity, bias, harm benchmarks.
- Don't trust a single benchmark: alignment is multi-dimensional.
- Keep human review: high-risk scenarios must have human-in-the-loop.