Benchmark
A Benchmark is a standard test set used to objectively measure model performance on a specific capability. Every serious LLM release publishes its scores on mainstream benchmarks.
Mainstream benchmarks by capability
Knowledge & Understanding
- MMLU: 57 subjects multiple choice, covering humanities/social sciences/STEM/law/medicine.
- C-Eval: Chinese comprehensive knowledge evaluation (Stanford + Tsinghua).
- GPQA: PhD-level science questions (Google-proof Q&A).
Reasoning & Code
- HumanEval: Python function completion (OpenAI).
- MBPP: basic Python programming problems.
- MATH: competition-grade math.
- SWE-bench (see swe-bench): real GitHub issue fixing.
Long Context
- RULER: 128k+ context comprehensive test.
- LongBench: bilingual long text.
- Needle in a Haystack: hide a needle in long text.
Agent Capability
- GAIA: multimodal, multi-step real-world Q&A.
- ToolBench: tool calling capability.
- WebArena: real browser-environment tasks.
Safety & Alignment
- TruthfulQA: factuality.
- HaluEval: hallucination detection.
Evaluation traps
- Data contamination: training/test overlap inflates scores.
- Overfitting benchmark: optimizing for one benchmark doesn't mean general ability.
- Prompt sensitivity: changing prompt template can swing scores by 10 points.
Practical advice: your own business-relevant, private eval set is more trustworthy than public benchmarks.