Skip to content

Artificial Analysis Intelligence Index what it measures, how it is scored and who leads it

Composite indexes · Artificial Analysis · Index

Scores from Artificial Analysis Intelligence Index, published Sep 27, 2026

RankModelIndex
1Claude Opus 5.5Anthropic · as “Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)”57.6
2Claude Fable 5.1Anthropic · as “Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)”53.4
3GPT-6 AstraOpenAI · as “GPT-6 Astra (max)”52.7
4Muse Spark 1.3Meta · as “Muse Spark 1.3 (max)”48.1
5GPT-6 SolOpenAI · as “GPT-6 Sol (max)”47.5
6Grok 4.7SpaceXAI · as “Grok 4.7 (xhigh)”46.4
7MiMo-V2.6-ProXiaomi · as “MiMo-V2.6-Pro”46.3
8Qwen3.8 Max (0902)Alibaba · as “Qwen3.8 Max (0902)”45.4
9GLM-5.3Z.ai · as “GLM-5.3 (max)”44.8
10Grok 4.6SpaceXAI · as “Grok 4.6 (high)”44.3
11Step 5 PreviewStepFun · as “Step 5 Preview”43.7
12Kimi K3Moonshot AI · as “Kimi K3 (max)”43.6
13GPT-5.6 TerraOpenAI · as “GPT-5.6 Terra (max)”42.1
14GLM 5.3 FlashZ.ai · as “GLM 5.3 Flash”41.8
15Gemini 3.8 FlashGoogle · as “Gemini 3.8 Flash (high)”40.9
16Qwen3.8 2.4T A95BAlibaba · as “Qwen3.8 2.4T A95B”39.9
17Qwen3.8-Flash-NextAlibaba · as “Qwen3.8-Flash-Next”39.8
18DeepSeek V4.1 FlashDeepSeek · as “DeepSeek V4.1 Flash (Reasoning, Max Effort)”39.5
19Claude Sonnet 5Anthropic · as “Claude Sonnet 5 (Adaptive Reasoning, Max Effort)”38.2
20GPT-6 LunaOpenAI · as “GPT-6 Luna (max)”37.3
21DeepSeek V4 Pro 0813DeepSeek · as “DeepSeek V4 Pro 0813 (Reasoning, Max Effort)”36.0
22DeepSeek V4 Flash VisionDeepSeek · as “DeepSeek V4 Flash Vision (Reasoning, Max Effort)”34.8
23Qwen3.8 27BAlibaba · as “Qwen3.8 27B (xhigh)”33.7
24Motif 3Motif Technologies · as “Motif 3”33.6
25GPT-5.3 CodexOpenAI · as “GPT-5.3 Codex (xhigh)”32.5

About this benchmark

Task
One number from ten independent evaluations that Artificial Analysis runs itself.
Dataset
Version 4.3.2 combines AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GDP.pdf and CritPt.
Method
Each model run by Artificial Analysis under the same settings; the index weights the evaluations and reports a 95% confidence interval.
Metric
Index
Organization
Artificial Analysis

Versions

Newest first. New versions are added, never rewritten.

VersionDate
4.3.2Current version as read on 2026-09-24 (ten evaluations).

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.