Skip to content

Benchmarks: what is being measured

85 AI benchmarks in 15 domains, each with its task, dataset, method, metric, versions and paper. 28 have scores captured here every day.

12 of 85 benchmarks in Coding

BenchmarkScores
HumanEvalCoding · 2021OpenAI·
MBPPCoding · 2021Google·
SWE-benchCoding · 2023Princeton University·
SWE-bench VerifiedCoding · 2024SWE-bench team and OpenAI39
SWE-Bench ProCoding · 2025Scale AI11
LiveCodeBenchCoding · 2024UC Berkeley, MIT and Cornell·
Aider PolyglotCoding · 2024Aider50
Terminal-BenchCoding · 2025Stanford University and Laude Institute·
SciCodeCoding · 2024University of Illinois and collaborators·
BigCodeBenchCoding · 2024BigCode·
MLE-benchCoding · 2024OpenAI·
SWE-LancerCoding · 2025OpenAI·

Scores counts the models with a captured score on the boards that publish this benchmark. A dot means we capture no board for it yet.

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.