Skip to content

SimpleQA what it measures, how it is scored and who leads it

Knowledge · OpenAI · introduced Nov 7, 2024 · Correct (and F-score)

About this benchmark

Task
Answer short fact-seeking questions with a single correct answer, or decline.
Dataset
Short questions collected adversarially against GPT-4 answers.
Method
Each answer graded correct, incorrect or not attempted by a model grader against the reference.
Metric
Correct (and F-score)
Organization
OpenAI
Introduced
Nov 7, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Nov 7, 20241 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.