Skip to content

MCP Atlas what it measures, how it is scored and who leads it

Agents and tools · Scale AI · Pass rate

Scores from MCP Atlas, published Sep 26, 2026

About this benchmark

Task
Complete realistic tasks by calling tools on real MCP servers, often across several servers.
Dataset
1,000 human-written tasks over 36 real MCP servers and 220 tools; the public leaderboard uses 500 of them.
Method
Three to six tool calls per task, with distractor tools exposed; about a third branch on earlier tool output.
Metric
Pass rate
Organization
Scale AI

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.