Skip to main content
ZICQ

Wiki Concepts

AI Alignment

Concepts
Aliases: alignment ·2026-09-14

AI Alignment

AI Alignment is the field of research that makes AI systems' behavior consistent with human intent, values, and expectations. It's both a core training objective for LLMs and the frontier of AI safety research.

Three layers of alignment

  1. Instruction Following: teach the model to understand and execute user intent.
  2. Value Alignment: make the model refuse harmful requests, output no dangerous content.
  3. Intent Alignment: make the model understand what the user really wants (not just literal).

Classic alignment stack

Pretraining → SFT → Reward Model → RLHF
                              ↘ DPO (critic-free simplified)
                              ↘ Constitutional AI (AI feedback)
  • SFT (sft): teach basic instruction following with (instruction, response) data.
  • RLHF (rlhf): train reward model from human preferences, optimize LLM with RL.
  • DPO (dpo): skip reward model, train directly on preference pairs.
  • Constitutional AI (Anthropic): use AI self-critique + revision for scalable alignment.

Manifestations of alignment failure

  • Hallucination: inventing facts (hallucination).
  • Deception: saying falsehoods to please users (sycophancy).
  • Jailbreak: bypassing safety constraints via prompt (prompt-injection).
  • Goal misspecification: optimizing the wrong target (reward hacking).
  • Power-seeking: model pursuing not being shut down (emergent, debated).

Frontier research directions

  • Scalable Oversight: extending weak supervision signals to superhuman capability (debate, recursive reward modeling).
  • Mechanistic Interpretability: understanding what models actually compute internally (Anthropic, DeepMind focus).
  • Constitutional AI / RLAIF: replacing expensive human labeling with AI.
  • Alignment Tax: minimizing the impact of alignment on general capability.

Practical recommendations

  • Production deployment must do alignment evaluation: toxicity, bias, harm benchmarks.
  • Don't trust a single benchmark: alignment is multi-dimensional.
  • Keep human review: high-risk scenarios must have human-in-the-loop.