Skip to content

Humanity's Last Exam what it measures, how it is scored and who leads it

Knowledge · Center for AI Safety and Scale AI · introduced Jan 24, 2025 · Accuracy

Scores from Humanity's Last Exam, published Sep 26, 2026

RankModelAccuracy
1GPT-6 AstraOpenAI · as “GPT 6 Astra”54.8%
2Claude Fable 5.1Anthropic · as “Fable 5.1 (xhigh)”46.5%
3Gemini 3.1 Pro PreviewGoogle · as “gemini-3.1-pro-preview (thinking high)”46.4%
4Gemini 3.8 FlashGoogle · as “Gemini 3.8 Flash”44.5%
5GPT-5.4 ProOpenAI · as “gpt-5.4-pro-2026-03-05”44.3%
6Muse SparkMeta · as “Muse Spark”40.6%
7Gemini 3 ProGoogle · as “gemini-3-pro-preview”37.5%
8GPT-5.4OpenAI · as “gpt-5.4-2026-03-05 (xhigh thinking)”36.2%
9Claude Opus 4.7Anthropic · as “claude-opus-4-7”36.2%
10Claude Opus 4.6Anthropic · as “claude-opus-4-6-thinking-max”34.4%
11GPT-5 ProOpenAI · as “gpt-5-pro-2025-10-06”31.6%
12GPT-5.2OpenAI · as “gpt-5.2-2025-12-11”27.8%
13GPT-5OpenAI · as “gpt-5-2025-08-07”25.3%
14Claude Opus 4.5Anthropic · as “claude-opus-4-5-20251101-thinking”25.2%
15Kimi K2.5Moonshot AI · as “kimi-k2.5”24.4%
16GPT-5.1OpenAI · as “gpt-5.1-thinking”23.7%
17Gemini 2.5 Pro (Jun 2025)Google · as “gemini-2.5-pro-preview-06-05”21.6%
18o3OpenAI · as “o3 (high) (April 2025)”20.3%
19GPT-5 miniOpenAI · as “gpt-5-mini-2025-08-07”19.4%
20o4-miniOpenAI · as “o4-mini (high) (April 2025)”18.1%
21Claude Sonnet 4.5Anthropic · as “claude-sonnet-4-5-20250929-thinking”13.7%
22Gemini 2.5 Flash (Sep 2025)Google · as “Gemini 2.5 Flash (April 2025)”12.1%
23Claude Opus 4.1Anthropic · as “claude-opus-4-1-20250805-thinking”11.5%
24Claude Opus 4Anthropic · as “Claude Opus 4 (Thinking)”10.7%
25Gemini 3.1 Flash-LiteGoogle · as “gemini-3.1-flash-lite-preview”8.6%

About this benchmark

Task
Answer expert-written closed-ended questions at the frontier of human knowledge, some with images.
Dataset
2,500 questions across dozens of subjects, including mathematics, humanities and the natural sciences.
Method
Exact-match and multiple-choice answers graded automatically; calibration error is reported next to accuracy.
Metric
Accuracy
Organization
Center for AI Safety and Scale AI
Introduced
Jan 24, 2025

Versions

Newest first. New versions are added, never rewritten.

VersionDate
HLE-RollingA continually updated fork from CAIS: cleaned questions, some easy ones replaced with harder held-out ones.Sep 17, 202610 days ago
FinalFinalised to 2,500 questions; earlier results moved to a legacy board.Apr 3, 20251 year ago
Initial releaseReleased with the paper.Jan 24, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.