Every AI leaderboard, side by side

Leaderboards28Models ranked835

Claude Fable 5.1 leads the AI leaderboards

28 leaderboards and 835 models so far (+27 in 30 days), captured daily · 15 language-model boards count, #1 worth 10 points · How it is scored

Latest moves

See all changes
Model
OpenRouter weekly usagenew version published (2026-10-04 to 2026-10-05) · Oct 6, 2026 today
GPT-5.6 SolOpenRouter weekly usage · score 1324.3 to 1297.3 (still #19) · Oct 6, 2026 today
GLM-5.2OpenRouter weekly usage · score 1331.1 to 1329.2 (still #18) · Oct 6, 2026 today
GPT-6 AstraOpenRouter weekly usage · score 1408.8 to 1388.2 (still #17) · Oct 6, 2026 today
Muse Spark 1.3 ContributorOpenRouter weekly usage · score 1433.4 to 1446.7 (still #16) · Oct 6, 2026 today
Kimi K3OpenRouter weekly usage · score 1619.3 to 1690.2 (still #15) · Oct 6, 2026 today
Hy3OpenRouter weekly usage · score 2071.6 to 1861.3 (still #14) · Oct 6, 2026 today
Gemini 3.8 FlashOpenRouter weekly usage · score 2182.2 to 2191.0 (still #13) · Oct 6, 2026 today
Claude Opus 5.5OpenRouter weekly usage · score 2463 to 2598.9 (still #12) · Oct 6, 2026 today
GLM-5.3OpenRouter weekly usage · score 2943.8 to 2970.3 (still #11) · Oct 6, 2026 today
Jev 1.13OpenRouter weekly usage · score 3100.1 to 3224.7 (still #10) · Oct 6, 2026 today
GPT-5.6 LunaOpenRouter weekly usage · score 4374.5 to 3366.0 (still #9) · Oct 6, 2026 today

Top 10 of 58 ranked models

Every model with specs, license and price

Compare:GPT vs ClaudeClaude vs GeminiAnthropic vs OpenAI vs GoogleOpen-weight vs closed

The podium wall

Language models: ranks on every board
ChatReasoningCodingAgents and toolsVision
Model Text ArenaMultiChallengeAA IntelligenceHLEGPQA DiamondEpoch ECIARC-AGI-2LiveBenchMMLU-ProWebDev ArenaSWE-Bench ProSWE-bench VerifiedAider PolyglotBFCLMCP AtlasSearch ArenaVision ArenaDocument Arena
1Claude Fable 5.1
2GPT-6 Astra
3Claude Opus 5.5
4Claude Fable 5
5GPT-6.1 Sol
6Claude Opus 5
7Claude Opus 4.6
8Claude Sonnet 5.5
#1 with the vendor’s logo23 podium7 top 10 stale board, not counted

More views

BoardScore
Text ArenaLMArena#1 Gemini 4 Argon1,525
MultiChallengeScale AI#1 Muse Spark75.5%
AA IntelligenceArtificial Analysis#1 Claude Opus 5.557.6
HLEScale AI and CAIS#1 GPT-6 Astra54.8%
GPQA DiamondEpoch AI#1 GPT-6 Astra95.8%
Epoch ECIEpoch AI#1 Claude Opus 5.5167.3
ARC-AGI-2ARC Prize#1 GPT-6 Astra95.0%
LiveBenchLiveBench#1 Claude Fable 5.183.4%
MMLU-ProTIGER-Lab#1 Gemini 3.1 Pro Preview91.2%stale
WebDev ArenaLMArena#1 Claude Opus 5.51,815
SWE-Bench ProScale AI#1 Claude Opus 598.0%
SWE-bench VerifiedSWE-bench#1 Claude Opus 4.576.8%stale
Aider PolyglotAider#1 GPT-588.0%stale
BFCLUC Berkeley Gorilla#1 Claude Opus 4.577.5%
MCP AtlasScale AI#1 Muse Spark 1.188.1%
Search ArenaLMArena#1 GPT-5.6 Sol1,257
Vision ArenaLMArena#1 Claude Fable 51,309
Document ArenaLMArena#1 Claude Opus 51,516
Text-to-ImageLMArena#1 GPT Image 2.5 Sunburst1,425
Image EditLMArena#1 GPT Image 2.5 Sunburst1,524
Text-to-VideoLMArena#1 Gemini Omni 1.1 Flash1,516
Image-to-VideoLMArena#1 MiniMax H31,495
Speech ArenaArtificial Analysis#1 Eleven V4 Turbo1,334
Open ASRHugging Face#1 scribe v2 pro (zoom)3.59 WER
MTEBMTEB#1 harrier-oss-v1-27b (microsoft)74.3%
AA SpeedArtificial Analysis#1 Gemini 3.8 Flash249 tok/s
AA PriceArtificial Analysis#1 MiMo-V2.6-Flash$0.17
OpenRouterOpenRouter#1 Space Bunny Alpha38.5T tok
All boards

Questions

Which AI model is best right now?

It depends on the job, which is why Models shows every major leaderboard side by side. The leader of leaders at the top of the home page is the model that finishes highest across all fresh language-model boards (chat, reasoning, coding, agents and vision), scored 10 points for #1 down to 1 point for #10 on each board.

What is the best LLM leaderboard?

There is no single one. LMArena measures what people prefer in blind votes, Artificial Analysis and Epoch AI run independent evaluation suites, Scale SEAL runs private expert benchmarks such as Humanity's Last Exam, SWE-bench and SWE-Bench Pro test real coding, BFCL and MCP Atlas test tool use. Models tracks them all and shows where they agree.

How is the leader of leaders calculated?

Each board is reduced to one entry per model (the best variant), ranked, and the top 10 earn 10, 9, 8 ... 1 points. Points are added across the boards counted, and the podium score is points earned over points available, from 0 to 100. Boards not updated in 180 days are shown but not counted. The full method is on the Method page.

How often is it updated?

Every day at 04:30 UTC, straight from each board's own page, data file or API. Each entry carries the date it was captured, and each board shows the date it last published.

Why do the same model names look different on each board?

Boards list variants: effort levels such as high or max, thinking modes, dated snapshots, and agent harnesses. Models maps them to one model and keeps the exact name each board published next to every entry.

Can I use the data?

Yes. /api/summary returns the computed standings as JSON, /llms-full.txt has every board's top 25 in plain text, and every number links back to the board that published it.