Skip to content

LongBench what it measures, how it is scored and who leads it

Long context · Tsinghua University and Zhipu AI · introduced Aug 28, 2023 · Average score

About this benchmark

Task
Understand long real documents in English and Chinese: QA, summarisation, code, few-shot.
Dataset
v1: 21 datasets in 6 categories. v2: 503 hard multiple-choice questions with contexts of 8,000 to 2 million words.
Method
Task metrics for v1; multiple-choice accuracy for v2.
Metric
Average score
Organization
Tsinghua University and Zhipu AI
Introduced
Aug 28, 2023
Official leaderboard
longbench2.github.io/

Versions

Newest first. New versions are added, never rewritten.

VersionDate
2LongBench v2 (arXiv 2412.15204).Dec 19, 20241 year ago
1Released with the paper.Aug 28, 20233 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.