MCP Atlas what it measures, how it is scored and who leads it
Agents and tools · Scale AI · Pass rate
Scores from MCP Atlas, published Sep 26, 2026
About this benchmark
- Task
- Complete realistic tasks by calling tools on real MCP servers, often across several servers.
- Dataset
- 1,000 human-written tasks over 36 real MCP servers and 220 tools; the public leaderboard uses 500 of them.
- Method
- Three to six tool calls per task, with distractor tools exposed; about a third branch on earlier tool output.
- Metric
- Pass rate
- Organization
- Scale AI
- Official leaderboard
- labs.scale.com/leaderboard/mcp_atlas
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.