RLHF (Reinforcement Learning from Human Feedback)
RLHF (Reinforcement Learning from Human Feedback) is the advanced form of fine-tuning: use human preference data to train a "reward model", then use RL (usually PPO) to maximize that reward signal in the LLM.
Three-step flow
- SFT: first do supervised fine-tuning to get a base model (see sft).
- Reward Model: collect human ranking data over multiple answers, train RM to learn "which answer is better".
- PPO training: use RM as reward signal, optimize LLM answers through RL.
Why RLHF is needed
- Instruction following: pure SFT tends to learn "format", not really understand "instructions".
- Safety: teach the model to refuse harmful requests.
- Style alignment: reward model encodes human preferences at finer granularity than SFT data.
Drawbacks
- Training instability: PPO hyperparameters are sensitive, prone to collapse.
- Reward hacking: model learns to "trick" the reward model (looks good, isn't useful).
- High cost: training RM + PPO is 3-5x more expensive than SFT.
Alternatives
- DPO (Direct Preference Optimization): skip RM and PPO, train directly on preference pairs. Simple, stable, near-RLHF performance, increasingly mainstream.
- RLAIF: use AI feedback instead of human feedback (Anthropic Constitutional AI).
- KTO / ORPO / SimPO: various DPO improvements.