Skip to main content
ZICQ

Wiki Concepts

Perplexity

Concepts
Aliases: PPL perplexity ·2026-09-14

Perplexity

Perplexity (PPL) is the most fundamental metric for evaluating language models. It measures how "surprised" a model is by a piece of text — the lower the perplexity, the more the model "understands" the text.

Mathematical definition

PPL(X) = exp(-1/N · Σ log p(x_i | x_<i))
  • N is the number of tokens
  • p(x_i | x_<i) is the model's probability of predicting the next token given prior context
  • Exponentiating converts log-likelihood to "equivalent branch count"

Intuition: PPL = 10 means the model on average predicts like choosing between 10 equal-probability options; PPL = 1 means perfect prediction.

Classic baselines

  • GPT-2: PPL ~29 (WikiText-103)
  • GPT-3: PPL ~20
  • Llama 3 8B: PPL ~6 (same task)
  • GPT-5 / Claude 4: PPL < 5

Every 5 PPL points down usually means significant quality improvement.

Use cases

  • Pretraining phase: monitor whether model is learning (run dev set PPL every 1000 steps).
  • Ablation experiments: compare PPL differences across architectures / training data / hyperparameters.
  • Cross-dataset comparison: PPL on fixed dataset is a fair metric.
  • Domain adaptation: evaluate whether model has reasonable PPL on vertical domains (medical / legal).

Scenarios where PPL doesn't apply

  • Cross-model: PPL is affected by tokenizer, can't directly compare models with different token counts.
  • Cross-dataset: PPL on different datasets isn't comparable.
  • Downstream tasks: low PPL doesn't mean good task performance (chatbot, QA, code generation need specific metrics).
  • Human alignment: PPL can't reflect helpful / harmless.

Practical experience

  • Watch PPL curve during training: steady decline + occasional spikes is normal; sustained no-decline is underfitting, sudden spike is explosion.
  • Compare PPL with same tokenizer: otherwise use bits-per-character (BPC) or other normalized metrics.
  • Small-sample PPL unstable: single text PPL fluctuates, need large test set (thousands of tokens) for average.
  • Surprisal debugging: single token's -log p(x) can debug "where the model is confused".

Alternative / complementary metrics

  • Bits Per Byte (BPB): comparable across tokenizers.
  • Cross-entropy Loss: one-to-one conversion with PPL, just without exp.
  • Downstream benchmarks (MMLU, HumanEval, etc.): task-oriented, closer to real value.