Skip to content

BIG-Bench Hard what it measures, how it is scored and who leads it

Reasoning · Google · introduced Oct 17, 2022 · Accuracy · saturated

About this benchmark

Task
Solve 23 BIG-Bench tasks where earlier models fell short of average human raters.
Dataset
23 tasks from BIG-Bench: logical deduction, date understanding, multistep arithmetic and more.
Method
Three-shot prompts with and without chain of thought; exact match.
Metric
Accuracy
Organization
Google
Introduced
Oct 17, 2022

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Oct 17, 20223 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.