Berkeley Function Calling Leaderboard (BFCL) what it measures, how it is scored and who leads it
Agents and tools · UC Berkeley (Gorilla) · introduced Feb 26, 2024 · Overall accuracy
Scores from Berkeley Function Calling Leaderboard, published Apr 12, 2026
About this benchmark
- Task
- Call functions and tools correctly: single, parallel, multi-turn, and agentic web search and memory.
- Dataset
- Expert-written and user-contributed function-calling cases across languages and API styles.
- Method
- Calls checked by abstract syntax tree matching and by executing them; v3 checks multi-turn state.
- Metric
- Overall accuracy
- Organization
- UC Berkeley (Gorilla)
- Introduced
- Feb 26, 2024
- Official leaderboard
- gorilla.cs.berkeley.edu/leaderboard.html
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| v4 AgenticWeb search, memory and prompt variation. | Jul 17, 20251 year ago | Web search, memory and prompt variation. |
| v3 Multi-turnMulti-turn and multi-step cases. | Sep 19, 20242 years ago | Multi-turn and multi-step cases. |
| v2 LiveUser-contributed live data. | Aug 14, 20242 years ago | User-contributed live data. |
| v1Launched. | Feb 26, 20242 years ago | Launched. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.