Humanity's Last Exam what it measures, how it is scored and who leads it
Knowledge · Center for AI Safety and Scale AI · introduced Jan 24, 2025 · Accuracy
Scores from Humanity's Last Exam, published Sep 26, 2026
About this benchmark
- Task
- Answer expert-written closed-ended questions at the frontier of human knowledge, some with images.
- Dataset
- 2,500 questions across dozens of subjects, including mathematics, humanities and the natural sciences.
- Method
- Exact-match and multiple-choice answers graded automatically; calibration error is reported next to accuracy.
- Metric
- Accuracy
- Organization
- Center for AI Safety and Scale AI
- Introduced
- Jan 24, 2025
- Official leaderboard
- labs.scale.com/leaderboard/humanitys_last_exam
- Home
- lastexam.ai
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| HLE-RollingA continually updated fork from CAIS: cleaned questions, some easy ones replaced with harder held-out ones. | Sep 17, 202610 days ago | A continually updated fork from CAIS: cleaned questions, some easy ones replaced with harder held-out ones. |
| FinalFinalised to 2,500 questions; earlier results moved to a legacy board. | Apr 3, 20251 year ago | Finalised to 2,500 questions; earlier results moved to a legacy board. |
| Initial releaseReleased with the paper. | Jan 24, 20251 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.