Skip to content

ARC-AGI-2 what it measures, how it is scored and who leads it

Reasoning · ARC Prize Foundation · introduced Mar 24, 2025 · Tasks solved

Scores from ARC-AGI-2, published Sep 24, 2026

RankModelTasks solved
1GPT-6 AstraOpenAI · as “GPT-6 Astra (Max)”95.0%
2Claude Opus 5.5Anthropic · as “Claude Opus 5.5 (High)”93.3%
3GPT-5.6 SolOpenAI · as “GPT-5.6 Sol (Max)”92.5%
4Claude Opus 5Anthropic · as “Claude Opus 5 (Max)”90.4%
5Claude Fable 5.1Anthropic · as “Claude Fable 5.1 (Max)”90.0%
6Claude Fable 5Anthropic · as “Claude Fable 5 (Max)”89.2%
6Gemini 3.8 FlashGoogle · as “Gemini 3.8 Flash (High)”89.2%
8GPT-5.5OpenAI · as “GPT-5.5 (XHigh)”85.0%
9Gemini 3.7 FlashGoogle · as “Gemini 3.7 Flash (High)”84.6%
10Gemini 3Google · as “Gemini 3 Deep Think (2/26)”84.6%
10GPT-5.5 ProOpenAI · as “GPT-5.5 Pro (High)”84.6%
12GPT-5.6 TerraOpenAI · as “GPT-5.6 Terra (Max)”83.9%
13GPT-5.4 ProOpenAI · as “GPT-5.4 Pro (XHigh)”83.3%
14Gemini 3.1 Pro PreviewGoogle · as “Gemini 3.1 Pro (Preview)”77.1%
15Dots3 NoteDots Studio · as “Dots3-Note Preview (Max)”76.8%
16Claude 4.7Anthropic · as “Claude 4.7 (Max)”75.8%
17GPT-5.4OpenAI · as “GPT-5.4 (XHigh)”74.0%
18Claude Opus 4.8Anthropic · as “Claude Opus 4.8 (High)”72.1%
18Gemini 3.5 FlashGoogle · as “Gemini 3.5 Flash (High)”72.1%
20Claude Opus 4.6Anthropic · as “Claude Opus 4.6 (120K, High)”69.2%
21Grok 4.6SpaceXAI · as “Grok 4.6 (XHigh)”67.1%
22Grok 4.20SpaceXAI · as “Grok 4.20 (Reasoning)”65.1%
23DeepSeek V4 Flash VisionDeepSeek · as “DeepSeek V4 Flash 0731 (Max)”61.4%
24DeepSeek V4 Pro 0813DeepSeek · as “DeepSeek V4 Pro 0813 (Max)”61.3%
25Claude Sonnet 4.6Anthropic · as “Claude Sonnet 4.6 (High)”60.4%

About this benchmark

Task
Solve harder abstract grid puzzles that are easy for people and hard for AI.
Dataset
1,000 public training tasks; public, semi-private and private evaluation sets of 120 tasks each, every task solved by at least two people.
Method
Two attempts per task (pass@2), cost per task reported next to the score.
Metric
Tasks solved
Organization
ARC Prize Foundation
Introduced
Mar 24, 2025
Official leaderboard
arcprize.org/leaderboard

Versions

Newest first. New versions are added, never rewritten.

VersionDate
Technical reportPaper with the human testing baseline.May 17, 20251 year ago
2Announced with ARC Prize 2025.Mar 24, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.