Skip to content

Benchmarks: what is being measured

85 AI benchmarks in 15 domains, each with its task, dataset, method, metric, versions and paper. 28 have scores captured here every day.

7 of 85 benchmarks in Safety and honesty

BenchmarkScores
HarmBenchSafety and honesty · 2024Center for AI Safety·
AgentHarmSafety and honesty · 2024UK AI Security Institute and Gray Swan AI·
StrongREJECTSafety and honesty · 2024UC Berkeley·
XSTestSafety and honesty · 2023Bocconi University and University of Oxford·
BBQSafety and honesty · 2021New York University·
WMDPSafety and honesty · 2024Center for AI Safety and Scale AI·
MASKSafety and honesty · 2025Center for AI Safety and Scale AI·

Scores counts the models with a captured score on the boards that publish this benchmark. A dot means we capture no board for it yet.

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.