Perplexity
Perplexity (PPL) is the most fundamental metric for evaluating language models. It measures how "surprised" a model is by a piece of text — the lower the perplexity, the more the model "understands" the text.
Mathematical definition
PPL(X) = exp(-1/N · Σ log p(x_i | x_<i))
- N is the number of tokens
- p(x_i | x_<i) is the model's probability of predicting the next token given prior context
- Exponentiating converts log-likelihood to "equivalent branch count"
Intuition: PPL = 10 means the model on average predicts like choosing between 10 equal-probability options; PPL = 1 means perfect prediction.
Classic baselines
- GPT-2: PPL ~29 (WikiText-103)
- GPT-3: PPL ~20
- Llama 3 8B: PPL ~6 (same task)
- GPT-5 / Claude 4: PPL < 5
Every 5 PPL points down usually means significant quality improvement.
Use cases
- Pretraining phase: monitor whether model is learning (run dev set PPL every 1000 steps).
- Ablation experiments: compare PPL differences across architectures / training data / hyperparameters.
- Cross-dataset comparison: PPL on fixed dataset is a fair metric.
- Domain adaptation: evaluate whether model has reasonable PPL on vertical domains (medical / legal).
Scenarios where PPL doesn't apply
- Cross-model: PPL is affected by tokenizer, can't directly compare models with different token counts.
- Cross-dataset: PPL on different datasets isn't comparable.
- Downstream tasks: low PPL doesn't mean good task performance (chatbot, QA, code generation need specific metrics).
- Human alignment: PPL can't reflect helpful / harmless.
Practical experience
- Watch PPL curve during training: steady decline + occasional spikes is normal; sustained no-decline is underfitting, sudden spike is explosion.
- Compare PPL with same tokenizer: otherwise use bits-per-character (BPC) or other normalized metrics.
- Small-sample PPL unstable: single text PPL fluctuates, need large test set (thousands of tokens) for average.
- Surprisal debugging: single token's -log p(x) can debug "where the model is confused".
Alternative / complementary metrics
- Bits Per Byte (BPB): comparable across tokenizers.
- Cross-entropy Loss: one-to-one conversion with PPL, just without exp.
- Downstream benchmarks (MMLU, HumanEval, etc.): task-oriented, closer to real value.