Open LLM Leaderboard what it measures, how it is scored and who leads it
Composite indexes · Hugging Face · introduced 2023 · Average
About this benchmark
- Task
- Rank open-weights models on a fixed set of academic benchmarks run the same way.
- Dataset
- v1: ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, GSM8K. v2: IFEval, BBH, MATH Level 5, GPQA, MuSR, MMLU-Pro.
- Method
- Every submitted model evaluated by Hugging Face with the same harness and settings.
- Metric
- Average
- Organization
- Hugging Face
- Introduced
- 2023
- Official leaderboard
- huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| RetiredResults stopped updating. | Mar 20, 20251 year ago | Results stopped updating. |
| v2Six harder benchmarks. | Jun 20242 years ago | Six harder benchmarks. |
| v1Six benchmarks. | 20233 years ago | Six benchmarks. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.