MATH (and MATH-500) what it measures, how it is scored and who leads it
Math · UC Berkeley · introduced Mar 5, 2021 · Accuracy · saturated
About this benchmark
- Task
- Solve competition mathematics problems and give the final answer.
- Dataset
- 12,500 competition problems in seven subjects and five difficulty levels. MATH-500 is a 500-problem test subset from OpenAI.
- Method
- Free-form answer, normalised and compared with the boxed reference answer.
- Metric
- Accuracy
- Organization
- UC Berkeley
- Introduced
- Mar 5, 2021
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| MATH-500OpenAI's 500-problem test subset, from Let's Verify Step by Step (arXiv 2305.20050). | May 31, 20233 years ago | OpenAI's 500-problem test subset, from Let's Verify Step by Step (arXiv 2305.20050). |
| 1Released with the paper. | Mar 5, 20215 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.