FrontierMath what it measures, how it is scored and who leads it
Math · Epoch AI · introduced Nov 7, 2024 · Problems solved
About this benchmark
- Task
- Solve original, unpublished research-level mathematics problems with automatically checkable answers.
- Dataset
- Hundreds of problems crafted and vetted by expert mathematicians; most are kept private.
- Method
- The model may run code; answers are checked automatically. Epoch runs the evaluations itself.
- Metric
- Problems solved
- Organization
- Epoch AI
- Introduced
- Nov 7, 2024
- Paper
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (arXiv 2411.04872)
- Official leaderboard
- epoch.ai/frontiermath
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Nov 7, 20241 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.