SimpleQA what it measures, how it is scored and who leads it
Knowledge · OpenAI · introduced Nov 7, 2024 · Correct (and F-score)
About this benchmark
- Task
- Answer short fact-seeking questions with a single correct answer, or decline.
- Dataset
- Short questions collected adversarially against GPT-4 answers.
- Method
- Each answer graded correct, incorrect or not attempted by a model grader against the reference.
- Metric
- Correct (and F-score)
- Organization
- OpenAI
- Introduced
- Nov 7, 2024
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Nov 7, 20241 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.