BIG-Bench Hard what it measures, how it is scored and who leads it
Reasoning · Google · introduced Oct 17, 2022 · Accuracy · saturated
About this benchmark
- Task
- Solve 23 BIG-Bench tasks where earlier models fell short of average human raters.
- Dataset
- 23 tasks from BIG-Bench: logical deduction, date understanding, multistep arithmetic and more.
- Method
- Three-shot prompts with and without chain of thought; exact match.
- Metric
- Accuracy
- Organization
- Introduced
- Oct 17, 2022
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Oct 17, 20223 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.