Skip to content

SWE-bench Verified what it measures, how it is scored and who leads it

Coding · SWE-bench team and OpenAI · introduced Aug 2024 · Resolved

Scores from SWE-bench Verified, published Feb 26, 2026

RankModelResolved
1Claude Opus 4.5Anthropic · as “Claude 4.5 Opus (high)”76.8%
2Gemini 3 FlashGoogle · as “Gemini 3 Flash (high)”75.8%
2MiniMax-M2.5MiniMax · as “MiniMax M2.5 (high)”75.8%
4Claude Opus 4.6Anthropic · as “Claude 4.6 Opus”75.6%
5Gemini 3 ProGoogle · as “Gemini 3 Pro Preview (2025-11-18)”74.2%
6GLM-5Z.ai · as “GLM 5 (high)”72.8%
6GPT-5.2OpenAI · as “GPT 5.2 (high)”72.8%
6GPT-5.2 CodexOpenAI · as “GPT 5.2 Codex”72.8%
9Claude Sonnet 4.5Anthropic · as “Claude 4.5 Sonnet (high)”71.4%
10Kimi K2.5Moonshot AI · as “Kimi K2.5 (high)”70.8%
11DeepSeek-V3.2DeepSeek · as “DeepSeek V3.2 (high)”70.0%
12Claude Opus 4Anthropic · as “Claude 4 Opus (20250514)”67.6%
13Claude 4.5 HaikuAnthropic · as “Claude 4.5 Haiku (high)”66.6%
14GPT-5.1OpenAI · as “GPT 5.1 (2025-11-13) (medium)”66.0%
14GPT 5.1 CodexOpenAI · as “GPT 5.1 Codex (medium)”66.0%
16GPT-5OpenAI · as “GPT 5 (2025-08-07) (medium)”65.0%
17Claude Sonnet 4Anthropic · as “Claude 4 Sonnet (20250514)”64.9%
18Kimi K2 ThinkingMoonshot AI · as “Kimi K2 Thinking”63.4%
19MiniMax M2MiniMax · as “MiniMax M2”61.0%
20GPT-5 miniOpenAI · as “GPT 5 mini (2025-08-07) (medium)”59.8%
21o3OpenAI · as “o3 (2025-04-16)”58.4%
22Devstral Small (2512)Mistral AI · as “Devstral Small (2512)”56.4%
23GLM 4.6Z.ai · as “GLM 4.6 (T=1)”55.4%
23Qwen3-Coder 480B/A35B InstructAlibaba · as “Qwen3-Coder 480B/A35B Instruct”55.4%
25GLM 4.5Z.ai · as “GLM 4.5 (2025-08-22)”54.2%

About this benchmark

Task
Resolve real GitHub issues; the subset people confirmed is well specified and fairly tested.
Dataset
500 SWE-bench tasks checked by professional software engineers.
Method
As SWE-bench. The board we capture runs every model in the same minimal bash-only agent, so models compare like for like.
Metric
Resolved
Organization
SWE-bench team and OpenAI
Introduced
Aug 2024
Official leaderboard
swebench.com/

Versions

Newest first. New versions are added, never rewritten.

VersionDate
VerifiedHuman-validated 500-task subset, with OpenAI.Aug 20242 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.