Skip to content

SWE-Bench Pro what it measures, how it is scored and who leads it

Coding · Scale AI · introduced Sep 21, 2025 · Resolved

Scores from SWE-Bench Pro, published Sep 27, 2026

About this benchmark

Task
Solve long-horizon, enterprise-style software engineering tasks in real repositories.
Dataset
1,865 problems from 41 actively maintained repositories: a public set, a held-out set and a commercial set from partner startups.
Method
Agents produce a patch graded by the repository tests. V2 locks the protocol: web tools off, every diff re-graded on a pristine image.
Metric
Resolved
Organization
Scale AI
Introduced
Sep 21, 2025

Versions

Newest first. New versions are added, never rewritten.

VersionDate
V2Refreshed public split of 642 tasks across 11 repositories (89 invalid tasks dropped), co-developed with Reflection.Sep 22, 20265 days ago
1Released with the paper; public set from 11 repositories.Sep 21, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.