MRCR (multi-round coreference) what it measures, how it is scored and who leads it
Long context · Google DeepMind; open version by OpenAI · introduced Sep 19, 2024 · Match ratio
About this benchmark
- Task
- Find the right one of several near-identical requests hidden in a long conversation and reproduce its answer.
- Dataset
- Synthetic long conversations with repeated similar requests.
- Method
- Similarity between the reproduced and reference answer, reported by context length.
- Metric
- Match ratio
- Organization
- Google DeepMind; open version by OpenAI
- Introduced
- Sep 19, 2024
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| Michelangelo MRCRIntroduced in the Michelangelo paper. | Sep 19, 20242 years ago | Introduced in the Michelangelo paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.