GDPval what it measures, how it is scored and who leads it
Agents and tools · OpenAI · introduced Oct 5, 2025 · Wins or ties against experts
About this benchmark
- Task
- Produce real work products (documents, slides, spreadsheets) for tasks from 44 occupations.
- Dataset
- Tasks written by professionals with about 14 years of experience, across the top 9 sectors of US GDP; a 220-task gold subset is open.
- Method
- Blind pairwise grading by industry experts against a human professional's deliverable; an automated grader for the gold set.
- Metric
- Wins or ties against experts
- Organization
- OpenAI
- Introduced
- Oct 5, 2025
- Paper
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (arXiv 2510.04374)
- Home
- evals.openai.com
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Oct 5, 202511 months ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.