Skip to content

MultiChallenge what it measures, how it is scored and who leads it

Chat and instructions · Scale AI · introduced Jan 29, 2025 · Average accuracy

Scores from Scale SEAL MultiChallenge, published Sep 26, 2026

RankModelScore
1Muse SparkMeta · as “Muse Spark”75.5%
2Muse Spark 1.1Meta · as “Muse Spark 1.1”75.3%
3Gemini 3.1 Pro PreviewGoogle · as “gemini-3.1-pro-preview”71.4%
4GPT-5.4 ProOpenAI · as “gpt-5.4-pro-2026-03-05”69.2%
5Gemini 3 ProGoogle · as “gemini-3-pro-preview”65.7%
6GPT-5.1OpenAI · as “gpt-5.1-2025-11-13-thinking”63.4%
7GPT-5OpenAI · as “gpt-5-thinking”63.2%
8o3-proOpenAI · as “o3-pro-2025-06-10-reasoning-high”62.4%
9Kimi K2.5Moonshot AI · as “kimi-k2.5”61.4%
10Gemini 3.1 Flash-LiteGoogle · as “gemini-3.1-flash-lite-preview”60.6%
11GPT-5 miniOpenAI · as “gpt-5-mini-thinking”59.0%
12Claude Opus 4.5Anthropic · as “claude-opus-4-5-20251101-thinking”59.0%
13Claude Opus 4Anthropic · as “claude-4-opus-thinking”58.6%
14Claude Opus 4.1Anthropic · as “claude-opus-4-1-20250805-thinking”57.2%
15Claude Sonnet 4Anthropic · as “claude-4-sonnet-thinking”57.1%
16o3OpenAI · as “o3-2025-04-16-reasoning-high”56.6%
17Claude Opus 4.6Anthropic · as “claude-opus-4-6 (Non-Thinking)”56.0%
18Kimi K2 ThinkingMoonshot AI · as “kimi-k2-thinking”55.4%
19Claude Sonnet 4.5Anthropic · as “claude-sonnet-4-5-20250929-thinking”55.3%
20Gemini 2.5 Pro (Jun 2025)Google · as “gemini-2-5-pro”53.6%
21Claude 3.7 SonnetAnthropic · as “claude-3-7-sonnet-thinking”51.6%
22GPT-5.1 InstantOpenAI · as “gpt-5.1-2025-11-13-instant”51.2%
23Claude 4.5 HaikuAnthropic · as “claude-haiku-4-5-20251001-thinking”50.5%
24DeepSeek V3P1DeepSeek · as “deepseek-v3p1”46.1%
25gpt-oss-120bOpenAI · as “gpt-oss-120b”45.3%

About this benchmark

Task
Hold realistic multi-turn conversations: retain instructions, remember context, stay self-coherent.
Dataset
Human-written multi-turn conversations in four challenge categories.
Method
Graded with instance-level rubrics by a model judge checked against people.
Metric
Average accuracy
Organization
Scale AI
Introduced
Jan 29, 2025

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Jan 29, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.