Skip to content

TruthfulQA what it measures, how it is scored and who leads it

Knowledge · University of Oxford and OpenAI · introduced Sep 8, 2021 · Truthful and informative rate · saturated

About this benchmark

Task
Answer questions where a common misconception tempts a false answer.
Dataset
Questions spanning health, law, finance and politics, written to elicit imitated falsehoods.
Method
Generation graded by humans or a fine-tuned judge; a multiple-choice variant scores the likelihood of true answers.
Metric
Truthful and informative rate
Organization
University of Oxford and OpenAI
Introduced
Sep 8, 2021

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Sep 8, 20215 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.