Skip to content

RULER what it measures, how it is scored and who leads it

Long context · NVIDIA · introduced Apr 9, 2024 · Accuracy by context length

About this benchmark

Task
Retrieve, trace and aggregate information across long synthetic contexts, beyond one needle.
Dataset
13 tasks with configurable sequence length.
Method
Synthetic tasks generated at each length; accuracy per length shows the real usable context.
Metric
Accuracy by context length
Organization
NVIDIA
Introduced
Apr 9, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Apr 9, 20242 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.