SWE-bench
SWE-bench (Software Engineering Benchmark) is the most authoritative code engineering benchmark: it asks models to fix real GitHub issues and submit patches that pass the project's own tests.
Differences from HumanEval
| Dimension | HumanEval | SWE-bench |
|---|---|---|
| Task | Complete standalone function | Fix cross-file issue |
| Context | Function signature + docstring | Entire repository |
| Evaluation | Single test passes | Full test suite |
| Realism | Tutorial problems | Real production code |
Datasets
- SWE-bench: 2,294 real Python issues from 12 open-source projects (Django, scikit-learn, matplotlib, astropy, pylint, etc.).
- SWE-bench Verified: 500 human-reviewed subset (OpenAI 2024), more reliable.
Mainstream model scores (2025)
| Model | SWE-bench Verified |
|---|---|
| Claude 4 Opus | ~72% |
| GPT-5 | ~65% |
| Claude 3.5 Sonnet | ~50% |
| DeepSeek V3 | ~42% |
Why it's the Agent benchmark
SWE-bench isn't a "fill in the function" problem, it's:
- Understand issue description (natural language)
- Locate relevant code in the repo
- Modify multiple files
- Write new tests
- Ensure existing tests still pass
A complete simulation of a human engineer's day. Hence the gold standard for Agent code ability.
Practical advice
If your project wants "AI auto-fix bugs / auto-PR", SWE-bench Verified scores are far more reliable than HumanEval.