LLM Reference
BFCL v3supersededAgentsTool use

BFCL v3: Berkeley Function Calling Leaderboard v3

Metric: Function Calling Accuracy (higher is better)Introduced: 2023Superseded by: bfcl

Version 3 of Berkeley Function Calling Leaderboard, evaluating model accuracy on function and API calls. The collected April 2026 slice has limited frontier-model coverage and is superseded by the current BFCL leaderboard.

Models ranked

14

tracked on this benchmark

Score band

73.7 – 49.1

best → lowest tracked

Snapshot trend

-5.30

Apr 12 → Jun 7 · 12 models

Leaderboard

Tracked models ranked by Function Calling Accuracy (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Granite 4.1 30B
Version: BFCL v3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
73.7
2
Qwen3.5-397B-A17B
Version: BFCL-V4, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
72.9
3
Qwen3.5-122B-A10B
Version: BFCL-V4, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
72.2
4
Qwen3-Max
Version: Berkeley Function Calling Leaderboard (BFCL v3)Harness: Not recordedEvaluator: Not recordedObserved: Apr 12, 2026Confidence: Not recordedSource
71.9
5
Qwen3-Coder-480B-A35B-Instruct
Version: Berkeley Function Calling Leaderboard (BFCL v3)Harness: Not recordedEvaluator: Not recordedObserved: Apr 12, 2026Confidence: Not recordedSource
68.7
6
Qwen3.5-27B
Version: BFCL-V4 (newest BFCL version), from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
68.5
7
Granite 4.1 8B
Version: BFCL v3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
68.3
8
Qwen3.5-35B-A3B
Version: BFCL-V4, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
67.3
9
Mellum2 12B
Version: BFCL v3 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
66.3
10
Qwen3.5-9B
Version: BFCL-V4, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
66.1
11
Kimi K2.5
Version: BFCL v3 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
64.5
12
Granite 4.1 3B
Version: BFCL v3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
60.8
13
Qwen3.5-4B
Version: BFCL-V4, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
50.3
14
LFM2.5 1.2B Instruct
Version: BFCLv3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
49.1

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: Apr 26, 2026

Resources