Skip to content

Open weights vs Closed weights

The best open-weight model on every board against the best closed one. One row per board where at least one is ranked, from each board's current capture.

Head to head on 28 boards: Closed weights places best on 25 and Open weights places best on 3. Both ranked on 28 of 28 boards.

BoardOpen weightsClosed weights
Chat2
Text ArenaArena ratingyesterday#131,488Kimi K3#11,509best of the pickedClaude Opus 5.5
MultiChallengeScoreyesterday#961.4%Kimi K2.5#175.5%best of the pickedMuse Spark
Reasoning7
AA IntelligenceIndexyesterday#746.3MiMo-V2.6-Pro#157.6best of the pickedClaude Opus 5.5
HLEAccuracyyesterday#1524.4%Kimi K2.5#154.8%best of the pickedGPT-6 Astra
GPQA DiamondAccuracyyesterday#1593.1%Kimi K3#195.8%best of the pickedGPT-6 Astra
Epoch ECIECIyesterday#12157.7Kimi K3#1166.6best of the pickedGPT-6 Astra
ARC-AGI-2Scoreyesterday#2361.4%DeepSeek V4 Flash Vision#195.0%best of the pickedGPT-6 Astra
LiveBenchGlobal averageyesterday#681.1%DeepSeek V4.1 Flash#183.4%best of the pickedClaude Fable 5.1
MMLU-ProAccuracy · staleyesterday#688.0%MiniMax M2.1#191.2%best of the pickedGemini 3.1 Pro Preview
Coding4
WebDev ArenaArena ratingyesterday#71,660Kimi K3#11,827best of the pickedClaude Opus 5.5
SWE-Bench ProResolvedyesterday#488.2%Kimi K3#198.0%best of the pickedClaude Opus 5
SWE-bench VerifiedResolved · staleyesterday#275.8%MiniMax-M2.5#176.8%best of the pickedClaude Opus 4.5
Aider PolyglotPass rate · staleyesterday#874.2%DeepSeek-V3.2#188.0%best of the pickedGPT-5
Agents and tools3
BFCLOverall accuracyyesterday#472.4%GLM 4.6#177.5%best of the pickedClaude Opus 4.5
MCP AtlasPass rateyesterday#484.5%Qwen3.8 2.4T A95B#188.1%best of the pickedMuse Spark 1.1
Search ArenaArena ratingyesterday#321,023Diffbot Small Xl#11,257best of the pickedGPT-5.6 Sol
Vision2
Vision ArenaArena ratingyesterday#241,275GLM 5.3 Flash#11,310best of the pickedClaude Fable 5
Document ArenaArena ratingyesterday#221,451Kimi K2.6#11,516best of the pickedClaude Opus 5
Image2
Text-to-ImageArena ratingyesterday#161,228Qwen Image 2.1#11,424best of the pickedGPT Image 2.5 Sunburst
Image EditArena ratingyesterday#151,367Qwen Image 2.1#11,526best of the pickedGPT Image 2.5 Sunburst
Video2
Text-to-VideoArena ratingyesterday#81,460MiniMax H3#11,516best of the pickedGemini Omni 1.1 Flash
Image-to-VideoArena ratingyesterday#11,495best of the pickedMiniMax H3#21,488Gemini Omni 1.1 Flash
Speech2
Speech ArenaArena ratingyesterday#91,206Breeze TTS 2#11,277best of the pickedSonic 3.6
Open ASRAverage WERyesterday#104.31 WERQwen3-ASR-1.7B-hf (Qwen)#13.59 WERbest of the pickedscribe v2 pro (zoom)
Embeddings1
MTEBMean task scoreyesterday#174.3%best of the pickedharrier-oss-v1-27b (microsoft)#568.4%gemini-embedding-001 (google)
Speed1
AA SpeedOutput speedyesterday#2233 tok/sDeepSeek V4.1 Flash#1328 tok/sbest of the pickedGemini 3.8 Flash
Price1
AA PriceBlended priceyesterday#2$0.23Qwen3.8-Flash-Next#1$0.20best of the pickedGPT-6 Luna
Usage1
OpenRouterWeekly tokensyesterday#119.2T tokbest of the pickedDeepSeek V4.1 Flash#48.6T tokGPT-5.6 Luna

Green marks the best placing among the picked where two or more are ranked; ties are marked for each. Ranks are each board's own order (lower is better), scores as the board published them. Dash: not on that board.

More comparisons

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.