LongBench what it measures, how it is scored and who leads it
Long context · Tsinghua University and Zhipu AI · introduced Aug 28, 2023 · Average score
About this benchmark
- Task
- Understand long real documents in English and Chinese: QA, summarisation, code, few-shot.
- Dataset
- v1: 21 datasets in 6 categories. v2: 503 hard multiple-choice questions with contexts of 8,000 to 2 million words.
- Method
- Task metrics for v1; multiple-choice accuracy for v2.
- Metric
- Average score
- Organization
- Tsinghua University and Zhipu AI
- Introduced
- Aug 28, 2023
- Official leaderboard
- longbench2.github.io/
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 2LongBench v2 (arXiv 2412.15204). | Dec 19, 20241 year ago | LongBench v2 (arXiv 2412.15204). |
| 1Released with the paper. | Aug 28, 20233 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.