SWE-bench Verified what it measures, how it is scored and who leads it
Coding · SWE-bench team and OpenAI · introduced Aug 2024 · Resolved
Scores from SWE-bench Verified, published Feb 26, 2026
About this benchmark
- Task
- Resolve real GitHub issues; the subset people confirmed is well specified and fairly tested.
- Dataset
- 500 SWE-bench tasks checked by professional software engineers.
- Method
- As SWE-bench. The board we capture runs every model in the same minimal bash-only agent, so models compare like for like.
- Metric
- Resolved
- Organization
- SWE-bench team and OpenAI
- Introduced
- Aug 2024
- Official leaderboard
- swebench.com/
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| VerifiedHuman-validated 500-task subset, with OpenAI. | Aug 20242 years ago | Human-validated 500-task subset, with OpenAI. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.