Skip to content

Open LLM Leaderboard what it measures, how it is scored and who leads it

Composite indexes · Hugging Face · introduced 2023 · Average

About this benchmark

Task
Rank open-weights models on a fixed set of academic benchmarks run the same way.
Dataset
v1: ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, GSM8K. v2: IFEval, BBH, MATH Level 5, GPQA, MuSR, MMLU-Pro.
Method
Every submitted model evaluated by Hugging Face with the same harness and settings.
Metric
Average
Organization
Hugging Face
Introduced
2023

Versions

Newest first. New versions are added, never rewritten.

VersionDate
RetiredResults stopped updating.Mar 20, 20251 year ago
v2Six harder benchmarks.Jun 20242 years ago
v1Six benchmarks.20233 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.