Skip to content

MATH (and MATH-500) what it measures, how it is scored and who leads it

Math · UC Berkeley · introduced Mar 5, 2021 · Accuracy · saturated

About this benchmark

Task
Solve competition mathematics problems and give the final answer.
Dataset
12,500 competition problems in seven subjects and five difficulty levels. MATH-500 is a 500-problem test subset from OpenAI.
Method
Free-form answer, normalised and compared with the boxed reference answer.
Metric
Accuracy
Organization
UC Berkeley
Introduced
Mar 5, 2021

Versions

Newest first. New versions are added, never rewritten.

VersionDate
MATH-500OpenAI's 500-problem test subset, from Let's Verify Step by Step (arXiv 2305.20050).May 31, 20233 years ago
1Released with the paper.Mar 5, 20215 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.