Skip to content

BIG-bench what it measures, how it is scored and who leads it

Reasoning · Google and 450+ collaborators · introduced Jun 9, 2022 · Normalised preferred metric

About this benchmark

Task
A broad collaborative suite of tasks believed to be beyond language models at the time.
Dataset
More than 200 tasks contributed by researchers across many institutions.
Method
Task-specific scoring, mostly exact match or multiple choice, normalised per task.
Metric
Normalised preferred metric
Organization
Google and 450+ collaborators
Introduced
Jun 9, 2022

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Jun 9, 20224 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.