MuSR what it measures, how it is scored and who leads it
Reasoning · University of Texas at Austin · introduced Oct 24, 2023 · Accuracy
About this benchmark
- Task
- Solve long natural-language murder mysteries, object placements and team allocations step by step.
- Dataset
- Algorithmically generated narratives with a hidden reasoning tree behind each question.
- Method
- Multiple choice after reading the story; chain of thought allowed.
- Metric
- Accuracy
- Organization
- University of Texas at Austin
- Introduced
- Oct 24, 2023
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Oct 24, 20232 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.