Self-Play
Self-play is a reinforcement-learning training paradigm: agents play against / collaborate with copies of themselves, using win/loss or feedback signals to keep evolving. It is the core method behind some of the strongest AI systems humans have built — from AlphaGo to AlphaZero, from AlphaProof to DeepSeek R1's GRPO, all lean heavily on this idea.
Key milestones
- TD-Gammon (1992, Gerald Tesauro): the first self-play RL system, reaching top human backgammon level
- AlphaGo (2016, DeepMind): self-play + human game records mixed training; 4:1 victory over Lee Sedol
- AlphaGo Zero (2017): pure self-play (no human game records) surpassed AlphaGo
- AlphaZero (2018): a general algorithm; Go / chess / shogi all from-scratch self-play
- OpenAI Five (2019): Dota 2, self-play + team coordination
- AlphaStar (2019): StarCraft II, self-play reached Grandmaster
- AlphaTensor (2022): self-play discovered new matrix-multiplication algorithms
- AlphaProof (2024): self-play for formal mathematical proofs
- DeepSeek R1 / GRPO (2024-2025): brought self-play thinking to LLM training (multiple samples from the same model + relative scoring)
How it works
Classic paradigm (games)
- Maintain an agent-snapshot pool
- Current agent plays against some older snapshot
- Win/loss becomes the reward
- Policy gradient updates the current agent
- Periodically push the current agent back into the pool
Benefits:
- Always has a "near-strength" opponent (no human data needed)
- Pool diversity avoids overfitting to the current self
LLM paradigm (GRPO / R1 style)
- Use the same model to generate K candidate answers for one prompt
- Use a verifier or relative scoring to judge which is better
- Convert "relative advantage" into reward
- PPO / GRPO training
This is RLVR + self-play combined:
- No human preference
- No external reward model (if a verifier is used)
- Agent compares with itself
Strengths
- No human data needed: pure self-play can reach superhuman levels from scratch
- Scalable: in theory, unlimited training time = unlimited progress
- Avoids reward hacking: opponents / opponent models are copies of itself, which evolve with the agent
- Cross-domain: board games / video games / math proofs / LLM reasoning all apply
Limitations
- Sparse rewards: win/loss is 0/1 feedback; training efficiency is bounded
- Strategy collapse: can fall into local optima (e.g. only knows one opening)
- Compute hungry: self-play demands massive large-model inference / game runs
- Out-of-distribution fragility: the "meta" of the training environment may not hold in deployment
- LLM self-play limits: pure self-play rarely learns new capabilities — typically combined with RLHF / RLVR / SFT