MultiChallenge what it measures, how it is scored and who leads it
Chat and instructions · Scale AI · introduced Jan 29, 2025 · Average accuracy
Scores from Scale SEAL MultiChallenge, published Sep 26, 2026
About this benchmark
- Task
- Hold realistic multi-turn conversations: retain instructions, remember context, stay self-coherent.
- Dataset
- Human-written multi-turn conversations in four challenge categories.
- Method
- Graded with instance-level rubrics by a model judge checked against people.
- Metric
- Average accuracy
- Organization
- Scale AI
- Introduced
- Jan 29, 2025
- Official leaderboard
- labs.scale.com/leaderboard/multichallenge
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Jan 29, 20251 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.