SWE-Bench Pro what it measures, how it is scored and who leads it
Coding · Scale AI · introduced Sep 21, 2025 · Resolved
Scores from SWE-Bench Pro, published Sep 27, 2026
About this benchmark
- Task
- Solve long-horizon, enterprise-style software engineering tasks in real repositories.
- Dataset
- 1,865 problems from 41 actively maintained repositories: a public set, a held-out set and a commercial set from partner startups.
- Method
- Agents produce a patch graded by the repository tests. V2 locks the protocol: web tools off, every diff re-graded on a pristine image.
- Metric
- Resolved
- Organization
- Scale AI
- Introduced
- Sep 21, 2025
- Official leaderboard
- labs.scale.com/leaderboard/swe_bench_pro_public_v2
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| V2Refreshed public split of 642 tasks across 11 repositories (89 invalid tasks dropped), co-developed with Reflection. | Sep 22, 20265 days ago | Refreshed public split of 642 tasks across 11 repositories (89 invalid tasks dropped), co-developed with Reflection. |
| 1Released with the paper; public set from 11 repositories. | Sep 21, 20251 year ago | Released with the paper; public set from 11 repositories. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.