BrowseComp what it measures, how it is scored and who leads it
Agents and tools · OpenAI · introduced Apr 16, 2025 · Accuracy
About this benchmark
- Task
- Find hard-to-locate, entangled facts by persistently browsing the web.
- Dataset
- 1,266 questions with short, verifiable answers.
- Method
- Short answer compared with the reference by a model grader.
- Metric
- Accuracy
- Organization
- OpenAI
- Introduced
- Apr 16, 2025
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Apr 16, 20251 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.