Every AI leaderboard, side by side

Leaderboards28Models ranked835

e5-small-v2 (intfloat) vs Muse Spark vs Qwen3.8 Max (0902) vs Kimi K3

One row per board where at least one is ranked, from each board's current capture.

  • e5-small-v2 (intfloat)
  • Muse Spark
  • Qwen3.8 Max (0902)
  • Kimi K3

Head to head on 10 boards: Qwen3.8 Max (0902) places best on 5, Kimi K3 places best on 4, Muse Spark places best on 1 and e5-small-v2 (intfloat) places best on 0. All 4 ranked on 0 of 17 boards.

Two at a time on this screen. Tap one to swap it in.

Boarde5-small-v2 (intfloat)Muse SparkQwen3.8 Max (0902)Kimi K3
Chat2
Text ArenaArena ratingyesterday#121,489best of the picked#201,482#131,488
MultiChallengeScoreyesterday#175.5%
Reasoning6
Epoch ECIECIyesterday#44152.0#22156.4#15157.4best of the picked
GPQA DiamondAccuracyyesterday#4189.8%#2192.7%#1893.1%best of the picked
AA IntelligenceIndexyesterday#1045.4best of the picked#1343.6
LiveBenchGlobal averageyesterday#1578.5%#1379.2%best of the picked
HLEAccuracyyesterday#640.6%
ARC-AGI-2Scoreyesterday#2960.4%
Coding2
WebDev ArenaArena ratingyesterday#91,671best of the picked#101,658
SWE-Bench ProResolvedyesterday#488.2%
Agents and tools1
MCP AtlasPass rateyesterday#982.2%#882.3%best of the picked
Vision2
Vision ArenaArena ratingyesterday#61,294#21,301best of the picked
Document ArenaArena ratingyesterday#241,444
Embeddings1
MTEBMean task scoreyesterday#6944.5%
Speed1
AA SpeedOutput speedyesterday#2439 tok/sbest of the picked#2536 tok/s
Price1
AA PriceBlended priceyesterday#14$3.00best of the picked#22$6.00
Usage1
OpenRouterWeekly tokensyesterday#151.7T tok

Green marks the best placing among the picked where two or more are ranked; ties are marked for each. Ranks are each board's own order (lower is better), scores as the board published them. Dash: not on that board.

More comparisons