TruthfulQA what it measures, how it is scored and who leads it
Knowledge · University of Oxford and OpenAI · introduced Sep 8, 2021 · Truthful and informative rate · saturated
About this benchmark
- Task
- Answer questions where a common misconception tempts a false answer.
- Dataset
- Questions spanning health, law, finance and politics, written to elicit imitated falsehoods.
- Method
- Generation graded by humans or a fine-tuned judge; a multiple-choice variant scores the likelihood of true answers.
- Metric
- Truthful and informative rate
- Organization
- University of Oxford and OpenAI
- Introduced
- Sep 8, 2021
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Sep 8, 20215 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.