Skip to main content
ZICQ

Wiki AI Concepts

Self-Play

AI Concepts
Aliases: Self-Play Self Play Self-Playing ·2026-09-19

Self-Play

Self-play is a reinforcement-learning training paradigm: agents play against / collaborate with copies of themselves, using win/loss or feedback signals to keep evolving. It is the core method behind some of the strongest AI systems humans have built — from AlphaGo to AlphaZero, from AlphaProof to DeepSeek R1's GRPO, all lean heavily on this idea.

Key milestones

  • TD-Gammon (1992, Gerald Tesauro): the first self-play RL system, reaching top human backgammon level
  • AlphaGo (2016, DeepMind): self-play + human game records mixed training; 4:1 victory over Lee Sedol
  • AlphaGo Zero (2017): pure self-play (no human game records) surpassed AlphaGo
  • AlphaZero (2018): a general algorithm; Go / chess / shogi all from-scratch self-play
  • OpenAI Five (2019): Dota 2, self-play + team coordination
  • AlphaStar (2019): StarCraft II, self-play reached Grandmaster
  • AlphaTensor (2022): self-play discovered new matrix-multiplication algorithms
  • AlphaProof (2024): self-play for formal mathematical proofs
  • DeepSeek R1 / GRPO (2024-2025): brought self-play thinking to LLM training (multiple samples from the same model + relative scoring)

How it works

Classic paradigm (games)

  1. Maintain an agent-snapshot pool
  2. Current agent plays against some older snapshot
  3. Win/loss becomes the reward
  4. Policy gradient updates the current agent
  5. Periodically push the current agent back into the pool

Benefits:

  • Always has a "near-strength" opponent (no human data needed)
  • Pool diversity avoids overfitting to the current self

LLM paradigm (GRPO / R1 style)

  1. Use the same model to generate K candidate answers for one prompt
  2. Use a verifier or relative scoring to judge which is better
  3. Convert "relative advantage" into reward
  4. PPO / GRPO training

This is RLVR + self-play combined:

  • No human preference
  • No external reward model (if a verifier is used)
  • Agent compares with itself

Strengths

  • No human data needed: pure self-play can reach superhuman levels from scratch
  • Scalable: in theory, unlimited training time = unlimited progress
  • Avoids reward hacking: opponents / opponent models are copies of itself, which evolve with the agent
  • Cross-domain: board games / video games / math proofs / LLM reasoning all apply

Limitations

  • Sparse rewards: win/loss is 0/1 feedback; training efficiency is bounded
  • Strategy collapse: can fall into local optima (e.g. only knows one opening)
  • Compute hungry: self-play demands massive large-model inference / game runs
  • Out-of-distribution fragility: the "meta" of the training environment may not hold in deployment
  • LLM self-play limits: pure self-play rarely learns new capabilities — typically combined with RLHF / RLVR / SFT