RULER what it measures, how it is scored and who leads it
Long context · NVIDIA · introduced Apr 9, 2024 · Accuracy by context length
About this benchmark
- Task
- Retrieve, trace and aggregate information across long synthetic contexts, beyond one needle.
- Dataset
- 13 tasks with configurable sequence length.
- Method
- Synthetic tasks generated at each length; accuracy per length shows the real usable context.
- Metric
- Accuracy by context length
- Organization
- NVIDIA
- Introduced
- Apr 9, 2024
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Apr 9, 20242 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.