Skip to content

Benchmarks: what is being measured

85 AI benchmarks in 15 domains, each with its task, dataset, method, metric, versions and paper. 28 have scores captured here every day.

85 of 85 benchmarks

BenchmarkScores
MMLUKnowledge · 2020UC Berkeley·
MMLU-ProKnowledge · 2024TIGER-Lab (University of Waterloo)100
GPQA (Diamond)Knowledge · 2023New York University100
Humanity's Last ExamKnowledge · 2025Center for AI Safety and Scale AI42
SimpleQAKnowledge · 2024OpenAI·
TruthfulQAKnowledge · 2021University of Oxford and OpenAI·
ARC-AGI-1Reasoning · 2019ARC Prize Foundation (created by François Chollet)·
ARC-AGI-2Reasoning · 2025ARC Prize Foundation79
ARC-AGI-3Reasoning · 2026ARC Prize Foundation·
BIG-Bench HardReasoning · 2022Google·
BIG-benchReasoning · 2022Google and 450+ collaborators·
HellaSwagReasoning · 2019University of Washington and Allen Institute for AI·
AI2 Reasoning Challenge (ARC)Reasoning · 2018Allen Institute for AI·
WinoGrandeReasoning · 2019Allen Institute for AI and University of Washington·
MuSRReasoning · 2023University of Texas at Austin·
LiveBenchReasoning · 2024LiveBench (Abacus.AI, NYU, NVIDIA and others)59
GSM8KMath · 2021OpenAI·
MATH (and MATH-500)Math · 2021UC Berkeley·
AIMEMathMathematical Association of America·
FrontierMathMath · 2024Epoch AI·
MGSMMath · 2022Google·
HumanEvalCoding · 2021OpenAI·
MBPPCoding · 2021Google·
SWE-benchCoding · 2023Princeton University·
SWE-bench VerifiedCoding · 2024SWE-bench team and OpenAI39
SWE-Bench ProCoding · 2025Scale AI11
LiveCodeBenchCoding · 2024UC Berkeley, MIT and Cornell·
Aider PolyglotCoding · 2024Aider50
Terminal-BenchCoding · 2025Stanford University and Laude Institute·
SciCodeCoding · 2024University of Illinois and collaborators·
BigCodeBenchCoding · 2024BigCode·
MLE-benchCoding · 2024OpenAI·
SWE-LancerCoding · 2025OpenAI·
τ-bench (tau-bench)Agents and tools · 2024Sierra·
Berkeley Function Calling Leaderboard (BFCL)Agents and tools · 2024UC Berkeley (Gorilla)81
MCP AtlasAgents and toolsScale AI34
GAIAAgents and tools · 2023Meta, Hugging Face and AutoGPT·
WebArenaAgents and tools · 2023Carnegie Mellon University·
OSWorldAgents and tools · 2024University of Hong Kong (XLang)·
BrowseCompAgents and tools · 2025OpenAI·
GDPvalAgents and tools · 2025OpenAI·
Vending-BenchAgents and tools · 2025Andon Labs·
IFEvalChat and instructions · 2023Google·
MultiChallengeChat and instructions · 2025Scale AI29
RULERLong context · 2024NVIDIA·
LongBenchLong context · 2023Tsinghua University and Zhipu AI·
∞Bench (InfiniteBench)Long context · 2024Tsinghua University and collaborators·
MRCR (multi-round coreference)Long context · 2024Google DeepMind; open version by OpenAI·
MMMUMultimodal · 2023MMMU team·
MathVistaMultimodal · 2023UCLA, University of Washington and Microsoft·
ChartQAMultimodal · 2022York University and Nanyang Technological University·
DocVQAMultimodal · 2020CVC Barcelona and IIIT Hyderabad·
Video-MMEMultimodal · 2024Video-MME team·
MMBenchMultimodal · 2023Shanghai AI Laboratory and collaborators·
AI2DMultimodal · 2016Allen Institute for AI·
GenEvalImage and video generation · 2023University of Washington·
VBenchImage and video generation · 2023Nanyang Technological University and Shanghai AI Laboratory·
Open ASR LeaderboardSpeechHugging Face and collaborators70
LibriSpeechSpeech · 2015Johns Hopkins University·
Common VoiceSpeech · 2019Mozilla·
FLEURSSpeech · 2022Google·
Artificial Analysis Speech ArenaSpeechArtificial Analysis90
MTEB (and MMTEB)Embeddings · 2022MTEB community (Hugging Face and collaborators)88
HarmBenchSafety and honesty · 2024Center for AI Safety·
AgentHarmSafety and honesty · 2024UK AI Security Institute and Gray Swan AI·
StrongREJECTSafety and honesty · 2024UC Berkeley·
XSTestSafety and honesty · 2023Bocconi University and University of Oxford·
BBQSafety and honesty · 2021New York University·
WMDPSafety and honesty · 2024Center for AI Safety and Scale AI·
MASKSafety and honesty · 2025Center for AI Safety and Scale AI·
LMArena Text Arena (Chatbot Arena)Arenas · 2023LMArena (from LMSYS, UC Berkeley)100
LMArena WebDev ArenaArenasLMArena100
LMArena Vision ArenaArenasLMArena100
LMArena Document ArenaArenasLMArena38
LMArena Search ArenaArenasLMArena33
LMArena Text-to-ImageArenasLMArena76
LMArena Image EditArenasLMArena53
LMArena Text-to-VideoArenasLMArena48
LMArena Image-to-VideoArenasLMArena48
Artificial Analysis Intelligence IndexComposite indexesArtificial Analysis100
Epoch Capabilities IndexComposite indexesEpoch AI100
Open LLM LeaderboardComposite indexes · 2023Hugging Face·
Artificial Analysis Output SpeedSpeed, price and usageArtificial Analysis24
Artificial Analysis PriceSpeed, price and usageArtificial Analysis24
OpenRouter usage rankingsSpeed, price and usageOpenRouter19

Scores counts the models with a captured score on the boards that publish this benchmark. A dot means we capture no board for it yet.

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.