Skip to content

∞Bench (InfiniteBench) what it measures, how it is scored and who leads it

Long context · Tsinghua University and collaborators · introduced Feb 21, 2024 · Average score

About this benchmark

Task
Answer questions over contexts longer than 100,000 tokens: books, code, math, dialogue.
Dataset
Tasks averaging over 100,000 tokens in English and Chinese.
Method
Task-specific exact match or multiple choice.
Metric
Average score
Organization
Tsinghua University and collaborators
Introduced
Feb 21, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Feb 21, 20242 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.