Skip to main content
ZICQ

Wiki Concepts

SWE-bench

Concepts
Aliases: SWE-bench Verified ·2026-09-14

SWE-bench

SWE-bench (Software Engineering Benchmark) is the most authoritative code engineering benchmark: it asks models to fix real GitHub issues and submit patches that pass the project's own tests.

Differences from HumanEval

Dimension HumanEval SWE-bench
Task Complete standalone function Fix cross-file issue
Context Function signature + docstring Entire repository
Evaluation Single test passes Full test suite
Realism Tutorial problems Real production code

Datasets

  • SWE-bench: 2,294 real Python issues from 12 open-source projects (Django, scikit-learn, matplotlib, astropy, pylint, etc.).
  • SWE-bench Verified: 500 human-reviewed subset (OpenAI 2024), more reliable.

Mainstream model scores (2025)

Model SWE-bench Verified
Claude 4 Opus ~72%
GPT-5 ~65%
Claude 3.5 Sonnet ~50%
DeepSeek V3 ~42%

Why it's the Agent benchmark

SWE-bench isn't a "fill in the function" problem, it's:

  • Understand issue description (natural language)
  • Locate relevant code in the repo
  • Modify multiple files
  • Write new tests
  • Ensure existing tests still pass

A complete simulation of a human engineer's day. Hence the gold standard for Agent code ability.

Practical advice

If your project wants "AI auto-fix bugs / auto-PR", SWE-bench Verified scores are far more reliable than HumanEval.