ARC-AGI-2 what it measures, how it is scored and who leads it
Reasoning · ARC Prize Foundation · introduced Mar 24, 2025 · Tasks solved
Scores from ARC-AGI-2, published Sep 24, 2026
About this benchmark
- Task
- Solve harder abstract grid puzzles that are easy for people and hard for AI.
- Dataset
- 1,000 public training tasks; public, semi-private and private evaluation sets of 120 tasks each, every task solved by at least two people.
- Method
- Two attempts per task (pass@2), cost per task reported next to the score.
- Metric
- Tasks solved
- Organization
- ARC Prize Foundation
- Introduced
- Mar 24, 2025
- Official leaderboard
- arcprize.org/leaderboard
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| Technical reportPaper with the human testing baseline. | May 17, 20251 year ago | Paper with the human testing baseline. |
| 2Announced with ARC Prize 2025. | Mar 24, 20251 year ago | Announced with ARC Prize 2025. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.