Skip to main content
ZICQ

Wiki Concepts

Benchmark

Concepts
Aliases: evaluation LLM benchmark ·2026-09-14

Benchmark

A Benchmark is a standard test set used to objectively measure model performance on a specific capability. Every serious LLM release publishes its scores on mainstream benchmarks.

Mainstream benchmarks by capability

Knowledge & Understanding

  • MMLU: 57 subjects multiple choice, covering humanities/social sciences/STEM/law/medicine.
  • C-Eval: Chinese comprehensive knowledge evaluation (Stanford + Tsinghua).
  • GPQA: PhD-level science questions (Google-proof Q&A).

Reasoning & Code

  • HumanEval: Python function completion (OpenAI).
  • MBPP: basic Python programming problems.
  • MATH: competition-grade math.
  • SWE-bench (see swe-bench): real GitHub issue fixing.

Long Context

  • RULER: 128k+ context comprehensive test.
  • LongBench: bilingual long text.
  • Needle in a Haystack: hide a needle in long text.

Agent Capability

  • GAIA: multimodal, multi-step real-world Q&A.
  • ToolBench: tool calling capability.
  • WebArena: real browser-environment tasks.

Safety & Alignment

  • TruthfulQA: factuality.
  • HaluEval: hallucination detection.

Evaluation traps

  • Data contamination: training/test overlap inflates scores.
  • Overfitting benchmark: optimizing for one benchmark doesn't mean general ability.
  • Prompt sensitivity: changing prompt template can swing scores by 10 points.

Practical advice: your own business-relevant, private eval set is more trustworthy than public benchmarks.