Skip to content

Benchmarks: what is being measured

85 AI benchmarks in 15 domains, each with its task, dataset, method, metric, versions and paper. 28 have scores captured here every day.

9 of 85 benchmarks in Agents and tools

BenchmarkScores
τ-bench (tau-bench)Agents and tools · 2024Sierra·
Berkeley Function Calling Leaderboard (BFCL)Agents and tools · 2024UC Berkeley (Gorilla)81
MCP AtlasAgents and toolsScale AI34
GAIAAgents and tools · 2023Meta, Hugging Face and AutoGPT·
WebArenaAgents and tools · 2023Carnegie Mellon University·
OSWorldAgents and tools · 2024University of Hong Kong (XLang)·
BrowseCompAgents and tools · 2025OpenAI·
GDPvalAgents and tools · 2025OpenAI·
Vending-BenchAgents and tools · 2025Andon Labs·

Scores counts the models with a captured score on the boards that publish this benchmark. A dot means we capture no board for it yet.

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.