Skip to main content
ZICQ

Wiki Concepts

RLHF

Concepts
Aliases: Reinforcement Learning from Human Feedback ·2026-09-14

RLHF (Reinforcement Learning from Human Feedback)

RLHF (Reinforcement Learning from Human Feedback) is the advanced form of fine-tuning: use human preference data to train a "reward model", then use RL (usually PPO) to maximize that reward signal in the LLM.

Three-step flow

  1. SFT: first do supervised fine-tuning to get a base model (see sft).
  2. Reward Model: collect human ranking data over multiple answers, train RM to learn "which answer is better".
  3. PPO training: use RM as reward signal, optimize LLM answers through RL.

Why RLHF is needed

  • Instruction following: pure SFT tends to learn "format", not really understand "instructions".
  • Safety: teach the model to refuse harmful requests.
  • Style alignment: reward model encodes human preferences at finer granularity than SFT data.

Drawbacks

  • Training instability: PPO hyperparameters are sensitive, prone to collapse.
  • Reward hacking: model learns to "trick" the reward model (looks good, isn't useful).
  • High cost: training RM + PPO is 3-5x more expensive than SFT.

Alternatives

  • DPO (Direct Preference Optimization): skip RM and PPO, train directly on preference pairs. Simple, stable, near-RLHF performance, increasingly mainstream.
  • RLAIF: use AI feedback instead of human feedback (Anthropic Constitutional AI).
  • KTO / ORPO / SimPO: various DPO improvements.