Scaling Laws
Scaling laws describe the power-law connection between large models' parameter count N, training tokens D, compute C, and test loss L. Empirically:
L(N, D) ≈ a / N^α + b / D^β + c
— make the model larger or train on more data, loss monotonically decreases at a predictable rate. This regularity is the core basis for designing the size of frontier models like GPT-3 / GPT-4 / Claude / Gemini / DeepSeek.
Key milestones
- 2020 Kaplan et al. (OpenAI): first systematic scaling laws, loss is a power law in N, D, C
- 2022 Hoffmann et al. (DeepMind): the Chinchilla paper proposes the "compute-optimal" model — given a compute budget, N and D should be scaled proportionally. 70B Chinchilla trained on 1.4T tokens beats same-compute Gopher (280B / 0.3T)
- 2023-2024: vendors follow Chinchilla to adjust training recipes (Llama 2 / Mistral)
- 2024-2025: in the RLHF / reasoning-model era, scaling laws extend to "inference compute" and "test-time compute"
Chinchilla's lesson
Compute-optimal, training tokens should ≈ 20× parameter count:
- 7B model should train on 140B tokens
- 70B model should train on 1.4T tokens
- 405B model should train on 8T+ tokens
This is why Llama 3 / Llama 4 push training data to 15T tokens.
Modern extensions
Test-time compute scaling (o1 / R1 paradigm)
- More compute at inference time (CoT / multiple samples) also buys performance
- Trades off with "make the model bigger":
- Big model: one-shot answer
- Small model + multi-step reasoning
- Both perform similarly at the same total compute
Emergent abilities
- Capabilities (multi-step reasoning, CoT, self-correction) that appear abruptly past some threshold
- Whether this is "true emergence" or a measurement artifact is academically contested (Schaeffer et al. 2023)
Diminishing returns
- Marginal benefit of data / parameters decreasing
- Data quality (FineWeb / RedPajama v2) > data quantity
- Current consensus: pure data / parameter scaling has hit a wall; RLHF / RLVR / test-time compute must take over
Engineering applications
- Predict training loss: extrapolate from small-model experiments to large-model final loss
- Resource allocation: compute-optimal trade-off between N / D / C
- When to stop training: extrapolate the scaling curve to determine "how much more data is worth it"
Limitations
- Architecture-dependent: transformer scaling laws don't transfer directly to Mamba / SSM
- Data quality dominates: pure web-data piling has hit the wall; synthetic data + curriculum learning matter more
- Diverse capability dimensions: loss dropping ≠ every task improves (reasoning, code, agent)
- Inference era breaks monotonic scale: test-time compute makes "scale up = scale down" possible