Skip to main content
ZICQ

Wiki AI Concepts

Scaling Laws

AI Concepts
Aliases: Scaling Laws Scaling Chinchilla Compute-Optimal ·2026-09-19

Scaling Laws

Scaling laws describe the power-law connection between large models' parameter count N, training tokens D, compute C, and test loss L. Empirically:

L(N, D) ≈ a / N^α + b / D^β + c

— make the model larger or train on more data, loss monotonically decreases at a predictable rate. This regularity is the core basis for designing the size of frontier models like GPT-3 / GPT-4 / Claude / Gemini / DeepSeek.

Key milestones

  • 2020 Kaplan et al. (OpenAI): first systematic scaling laws, loss is a power law in N, D, C
  • 2022 Hoffmann et al. (DeepMind): the Chinchilla paper proposes the "compute-optimal" model — given a compute budget, N and D should be scaled proportionally. 70B Chinchilla trained on 1.4T tokens beats same-compute Gopher (280B / 0.3T)
  • 2023-2024: vendors follow Chinchilla to adjust training recipes (Llama 2 / Mistral)
  • 2024-2025: in the RLHF / reasoning-model era, scaling laws extend to "inference compute" and "test-time compute"

Chinchilla's lesson

Compute-optimal, training tokens should ≈ 20× parameter count:

  • 7B model should train on 140B tokens
  • 70B model should train on 1.4T tokens
  • 405B model should train on 8T+ tokens

This is why Llama 3 / Llama 4 push training data to 15T tokens.

Modern extensions

Test-time compute scaling (o1 / R1 paradigm)

  • More compute at inference time (CoT / multiple samples) also buys performance
  • Trades off with "make the model bigger":
    • Big model: one-shot answer
    • Small model + multi-step reasoning
    • Both perform similarly at the same total compute

Emergent abilities

  • Capabilities (multi-step reasoning, CoT, self-correction) that appear abruptly past some threshold
  • Whether this is "true emergence" or a measurement artifact is academically contested (Schaeffer et al. 2023)

Diminishing returns

  • Marginal benefit of data / parameters decreasing
  • Data quality (FineWeb / RedPajama v2) > data quantity
  • Current consensus: pure data / parameter scaling has hit a wall; RLHF / RLVR / test-time compute must take over

Engineering applications

  • Predict training loss: extrapolate from small-model experiments to large-model final loss
  • Resource allocation: compute-optimal trade-off between N / D / C
  • When to stop training: extrapolate the scaling curve to determine "how much more data is worth it"

Limitations

  • Architecture-dependent: transformer scaling laws don't transfer directly to Mamba / SSM
  • Data quality dominates: pure web-data piling has hit the wall; synthetic data + curriculum learning matter more
  • Diverse capability dimensions: loss dropping ≠ every task improves (reasoning, code, agent)
  • Inference era breaks monotonic scale: test-time compute makes "scale up = scale down" possible