Skip to content

HellaSwag what it measures, how it is scored and who leads it

Reasoning · University of Washington and Allen Institute for AI · introduced May 19, 2019 · Accuracy · saturated

About this benchmark

Task
Pick the most plausible continuation of an everyday scenario.
Dataset
Four-option sentence completions with adversarially filtered wrong endings.
Method
Multiple choice, usually by comparing the likelihood of each ending.
Metric
Accuracy
Organization
University of Washington and Allen Institute for AI
Introduced
May 19, 2019

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.May 19, 20197 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.