LLM Reference

The tool use leaderboard · for developers

Best for tool use

Decision area
Reliable function calling and structured outputs.
Editor picks
6 editor picks
Eligible models
12 eligible models
See raw /best
EDITOR'S CHOICEResearched 8d ago

Claude Sonnet 5

Anthropic · 1m context
Excellent

Structured output is table stakes — pick on schema reliability and price.

Current Sonnet for function calling and structured outputs: 1M window, 86.6% BrowseComp multi-agent, 81.2% OSWorld-Verified, list $3 / $15 per 1M tokens. No BFCL row on the live model page.

The numbers
$/1M out
$10.00
$2.00 input
Context
1m
max window
Pros
  • +86.6% BrowseComp multi-agent
  • +81.2% OSWorld-Verified
  • +1M context at Sonnet list price
Cons
  • No BFCL row on the live model page

Also worth picking

The runners-up

ranked by editorial pick orderEditorial tiersExcellentStrongSolid
Google DeepMind · 1m
$5.00 / 1M out
Best current BFCL (72.5) with rock-solid JSON-schema adherence and a 1M window at $5 out.
Anthropic · 1m
$25.00 / 1M out
Editorial candidate with native tool use, structured outputs, mid-conversation tool changes, and a 1M window. Keep it off a tool-use podium claim until a compatible BFCL or τ-bench row is available.
Anthropic · 1m
$15.00 / 1M out
Best at picking the right tool when ten look plausible; pairs schema discipline with τ-bench leadership.
OpenAI · 1.05m
$15.00 / 1M out
Matches GPT-5.5's structured-output reliability at half the output price ($15 vs $30) — the GPT to run at high call volume.
Alibaba · 262k
$2.34 / 1M out
BFCL 72.9 with open weights — strong function calling you can self-host.

Eligibility

12 models are eligible for this board

Eligibility means tagged with useCases: [tool-use]. Pins must come from this pool.
All picks