Skip to content

Aider Polyglot what it measures, how it is scored and who leads it

Coding · Aider · introduced Dec 2024 · Pass rate

Scores from Aider Polyglot, published Oct 3, 2025

RankModelPass rate
1GPT-5OpenAI · as “gpt-5 (high)”88.0%
2o3-proOpenAI · as “o3-pro (high)”84.9%
3Gemini 2.5 Pro (Jun 2025)Google · as “gemini-2.5-pro-preview-06-05 (32k think)”83.1%
4o3OpenAI · as “o3 (high)”81.3%
5Grok 4SpaceXAI · as “grok-4 (high)”79.6%
6o3 GPT 4.1OpenAI · as “o3 (high) + gpt-4.1”78.2%
7Gemini 2.5 Pro 05 06Google · as “Gemini 2.5 Pro Preview 05-06”76.9%
8DeepSeek-V3.2DeepSeek · as “DeepSeek-V3.2-Exp (Reasoner)”74.2%
9Gemini 2.5 Pro 03 25Google · as “Gemini 2.5 Pro Preview 03-25”72.9%
10Claude Opus 4Anthropic · as “claude-opus-4-20250514 (32k thinking)”72.0%
10o4-miniOpenAI · as “o4-mini (high)”72.0%
12DeepSeek-R1 (May 2025)DeepSeek · as “DeepSeek R1 (0528)”71.4%
13Claude 3.7 SonnetAnthropic · as “claude-3-7-sonnet-20250219 (32k thinking tokens)”64.9%
14DeepSeek R1 Claude 3 5 SonnetDeepSeek · as “DeepSeek R1 + claude-3-5-sonnet-20241022”64.0%
15o1OpenAI · as “o1-2024-12-17 (high)”61.7%
16Claude Sonnet 4Anthropic · as “claude-sonnet-4-20250514 (32k thinking)”61.3%
17o3 MiniOpenAI · as “o3-mini (high)”60.4%
18Qwen3 235b A22B Diff No AlibabaAlibaba · as “Qwen3 235B A22B diff, no think, Alibaba API”59.6%
19Kimi K2 ThinkingMoonshot AI · as “Kimi K2”59.1%
20DeepSeek V3DeepSeek · as “DeepSeek V3 (0324)”55.1%
20Gemini 2.5 Flash (Sep 2025)Google · as “gemini-2.5-flash-preview-05-20 (24k think)”55.1%
22Quasar Alpha · as “Quasar Alpha”54.7%
23Grok 3SpaceXAI · as “Grok 3 Beta”53.3%
24Optimus Alpha · as “Optimus Alpha”52.9%
25GPT 4.1OpenAI · as “gpt-4.1”52.4%

About this benchmark

Task
Edit code to solve hard programming exercises in six languages, through the Aider coding assistant.
Dataset
225 of the hardest Exercism exercises in C++, Go, Java, JavaScript, Python and Rust.
Method
Two tries per exercise; the second sees the failing test output. Edits must apply cleanly in Aider's edit format.
Metric
Pass rate
Organization
Aider
Introduced
Dec 2024
Official leaderboard
aider.chat/docs/leaderboards/

Versions

Newest first. New versions are added, never rewritten.

VersionDate
Polyglot225 exercises in six languages.Dec 20241 year ago
Code editing (Python)The earlier Aider benchmark on Python Exercism exercises.20233 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.