∞Bench (InfiniteBench) what it measures, how it is scored and who leads it
Long context · Tsinghua University and collaborators · introduced Feb 21, 2024 · Average score
About this benchmark
- Task
- Answer questions over contexts longer than 100,000 tokens: books, code, math, dialogue.
- Dataset
- Tasks averaging over 100,000 tokens in English and Chinese.
- Method
- Task-specific exact match or multiple choice.
- Metric
- Average score
- Organization
- Tsinghua University and collaborators
- Introduced
- Feb 21, 2024
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Feb 21, 20242 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.