Skip to content

Berkeley Function Calling Leaderboard (BFCL) what it measures, how it is scored and who leads it

Agents and tools · UC Berkeley (Gorilla) · introduced Feb 26, 2024 · Overall accuracy

Scores from Berkeley Function Calling Leaderboard, published Apr 12, 2026

RankModelScore
1Claude Opus 4.5Anthropic · as “Claude-Opus-4-5-20251101 (FC)”77.5%
2Claude Sonnet 4.5Anthropic · as “Claude-Sonnet-4-5-20250929 (FC)”73.2%
3Gemini 3 ProGoogle · as “Gemini-3-Pro-Preview (Prompt)”72.5%
4GLM 4.6Z.ai · as “GLM-4.6 (FC thinking)”72.4%
5Grok 4 1 FastSpaceXAI · as “Grok-4-1-fast-reasoning (FC)”69.6%
6Claude 4.5 HaikuAnthropic · as “Claude-Haiku-4-5-20251001 (FC)”68.7%
7o3OpenAI · as “o3-2025-04-16 (Prompt)”63.0%
8Grok 4SpaceXAI · as “Grok-4-0709 (Prompt)”63.0%
9Moonshotai Kimi K2Moonshot AI · as “Moonshotai-Kimi-K2-Instruct (FC)”59.1%
10Grok 4 1 Fast NonSpaceXAI · as “Grok-4-1-fast-non-reasoning (FC)”58.3%
11Command A+Cohere · as “Command A Reasoning (FC)”57.1%
12DeepSeek-V3.2DeepSeek · as “DeepSeek-V3.2-Exp (Prompt + Thinking)”56.7%
13Gemini 2.5 Flash (Sep 2025)Google · as “Gemini-2.5-Flash (FC)”56.2%
14GPT-5.2OpenAI · as “GPT-5.2-2025-12-11 (FC)”55.9%
15GPT-5 miniOpenAI · as “GPT-5-mini-2025-08-07 (FC)”55.5%
16Xlam 2 32b Fc RSalesforce · as “xLAM-2-32b-fc-r (FC)”54.7%
17GPT 4.1OpenAI · as “GPT-4.1-2025-04-14 (FC)”54.0%
18o4-miniOpenAI · as “o4-mini-2025-04-16 (FC)”53.2%
19Xlam 2 70b Fc RSalesforce · as “xLAM-2-70b-fc-r (FC)”53.1%
20Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba · as “Qwen3-235B-A22B-Instruct-2507 (Prompt)”52.1%
21GPT 5 nanoOpenAI · as “GPT-5-nano-2025-08-07 (FC)”51.5%
22Nanbeige4 3bNanbeige · as “Nanbeige4-3B-Thinking-2511 (FC)”51.4%
23GPT 4.1 miniOpenAI · as “GPT-4.1-mini-2025-04-14 (FC)”50.5%
24Qwen3 32bAlibaba · as “Qwen3-32B (FC)”48.7%
25Nanbeige3.5 ProNanbeige · as “Nanbeige3.5-Pro-Thinking (FC)”47.7%

About this benchmark

Task
Call functions and tools correctly: single, parallel, multi-turn, and agentic web search and memory.
Dataset
Expert-written and user-contributed function-calling cases across languages and API styles.
Method
Calls checked by abstract syntax tree matching and by executing them; v3 checks multi-turn state.
Metric
Overall accuracy
Organization
UC Berkeley (Gorilla)
Introduced
Feb 26, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
v4 AgenticWeb search, memory and prompt variation.Jul 17, 20251 year ago
v3 Multi-turnMulti-turn and multi-step cases.Sep 19, 20242 years ago
v2 LiveUser-contributed live data.Aug 14, 20242 years ago
v1Launched.Feb 26, 20242 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.