Skip to content

MuSR what it measures, how it is scored and who leads it

Reasoning · University of Texas at Austin · introduced Oct 24, 2023 · Accuracy

About this benchmark

Task
Solve long natural-language murder mysteries, object placements and team allocations step by step.
Dataset
Algorithmically generated narratives with a hidden reasoning tree behind each question.
Method
Multiple choice after reading the story; chain of thought allowed.
Metric
Accuracy
Organization
University of Texas at Austin
Introduced
Oct 24, 2023

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Oct 24, 20232 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.