SWE-bench what it measures, how it is scored and who leads it
Coding · Princeton University · introduced Oct 10, 2023 · Resolved
About this benchmark
- Task
- Resolve a real GitHub issue by editing the repository so the hidden tests pass.
- Dataset
- 2,294 issues and pull requests from 12 popular Python repositories.
- Method
- The model (usually inside an agent) produces a patch; it counts if the tests that the real fix made pass now pass and nothing else breaks.
- Metric
- Resolved
- Organization
- Princeton University
- Introduced
- Oct 10, 2023
- Official leaderboard
- swebench.com/
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| MultimodalVisual JavaScript issues (arXiv 2410.03859). | Oct 4, 20241 year ago | Visual JavaScript issues (arXiv 2410.03859). |
| ContainerisedDocker images for reproducible evaluation. | Jun 20242 years ago | Docker images for reproducible evaluation. |
| LiteA smaller, cheaper subset. | Mar 20242 years ago | A smaller, cheaper subset. |
| Full2,294 tasks, released with the paper. | Oct 10, 20232 years ago | 2,294 tasks, released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.