Best LLMs for JSON / Tool use (2026)

Last refreshed 2026-09-29. Next refresh: weekly.

Best LLMs for JSON / Tool use in 2026. Ranked by BFCL benchmark, with native JSON output and structured output support.

Verdict

Use Gemini 3 Pro for JSON / Tool use today.

GPT-6.1 Sol is the runner-up; compare BFCL against Release.

Researched 8d agoWhy this pickMethodology

How we rank

JSON / Tool use leaders rank on BFCL first; when Berkeley coverage lags new SKUs, we fall back to τ-bench so fresh agentic models still surface.

  1. Eligibility — Models with `function_calling` or `tool_use` enabled in seed metadata.
  2. Primary ranking — BFCL score if present, otherwise τ-bench, then newer release.
  3. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. Pricing — Lowest tracked input(output) for API routes.

New models awaiting benchmark coverage

These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.
ModelWhy it is listedStatusTracked price
GPT-6.1 Sol
ToolsCode execution
GPT-6.1 Sol reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land.Benchmark pending

No tracked BFCL or tau-bench score yet.

In $2.00 / Out $10.00
Claude Sonnet 5.5
Tools
Claude Sonnet 5.5 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land.Benchmark pending

No tracked BFCL or tau-bench score yet.

In $2.00 / Out $10.00
Perceptron Mk1.5
Tools
Perceptron Mk1.5 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land.Benchmark pending

No tracked BFCL or tau-bench score yet.

In $0.15 / Out $1.50
Ember-1
Tools
Ember-1 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land.Benchmark pending

No tracked BFCL or tau-bench score yet.

In $3.00 / Out $15.00
#ModelInput $/1MOutput $/1M
1Ring-2.6-1T
ReasoningTools

Signal used: τ-bench 95.32%

$0.07$0.63
2Mistral Medium 3.5
ReasoningVisionTools

Signal used: τ-bench 91.4%

$1.50$7.50
3ByteDance Doubao Seed 2.0 Pro
VisionTools

Signal used: τ-bench 90.4%

$0.47$2.37
4Claude Mythos Preview
Invite-onlyReasoningVisionTools

Signal used: τ-bench 89.2%

——
5LFM2.5 8B A1B
ReasoningTools

Signal used: τ-bench 88.07%

——
6Claude Sonnet 4.6
ReasoningVisionTools

Signal used: τ-bench 87.5%

$3.00$15.00
7Command A+
ReasoningVisionTools

Signal used: τ-bench 85%

——
8Claude Opus 4.6
ReasoningVisionTools

Signal used: τ-bench 84.8%

$5.00$25.00
9GLM-5
ReasoningTools

Signal used: τ-bench 82.1%

$0.60$2.08
10Qwen3.5-35B-A3B
ReasoningTools

Signal used: τ-bench 81.2%

$0.14$1.00
11Qwen3.5-122B-A10B
ReasoningVisionTools

Signal used: τ-bench 79.5%

$0.26$2.08
12Qwen3.5-9B
VisionTools

Signal used: τ-bench 79.1%

$0.10$0.15
13Qwen3.5-27B
ReasoningVisionTools

Signal used: τ-bench 79%

$0.20$1.56
14Grok 4.20
ReasoningVisionTools

Signal used: τ-bench 78.9%

$1.25$2.50
15GPT-5.4
ReasoningVisionTools

Signal used: τ-bench 78.3%

$2.50$15.00
16GPT-5.3-Codex
ReasoningVisionTools

Signal used: τ-bench 77.8%

$1.75$14.00
17Claude Opus 4.5
ReasoningVisionTools

Signal used: BFCL 77.47%

$5.00$25.00
18Qwen3.6-Plus
VisionTools

Signal used: τ-bench 76.8%

$0.33$1.95
19Qwen3-Max
VisionTools

Signal used: τ-bench 76.8%

$0.78$3.90
20Gemini 3.1 Pro Preview
PreviewVisionTools

Signal used: τ-bench 76.5%

$2.00$12.00

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • MiniMax M3.1 Flash Preview (API id MiniMax-M3.1-Flash-Preview) is MiniMax's faster/lighter M3-series multimodal coding model announced for MiniMax Code on 2026-09-27 and made live on Token Plan the same day. First-party docs: 1,000,000-token context; text+image+video inputs to text; agentic reasoning, tool use, and coding; thinking always on with tunable effort low|medium|high|xhigh|max (default max; cannot disable — 400 'requires adaptive thinking'). Availability for now is Token Plan + MiniMax Code only — NOT on pay-as-you-go (paygo price table lists M3/M2.x only; no M3.1 Flash $/MTok row). No public Hugging Face weights/model card for the Flash Preview SKU. Closed-source hosted API — licenseSlug proprietary / weights notReleased.

    2026-09-27

    Release

  • Perceptron Mk1.5 is Perceptron's embodied vision-language model released September 25, 2026 (first-party blog). Successor to Mk1 in the Perceptron Mk family: native audio in, video object tracking, web search, sub-agent/tool calls, and stronger visual reasoning; blog cites ~2-5x faster end-to-end vs Mk1. Accepts text, images, video, and audio; emits text plus spatial/temporal annotations (points, boxes, polygons, clips, object tracks). First-party docs: Model ID perceptron-mk1.5; context window 36,864 tokens; max output 8,192; pricing $0.15/M input, $1.50/M output, $0.0375/M cached input. Available via Perceptron Platform + SDK and OpenRouter (perceptron/perceptron-mk1.5). Closed-source hosted API under a proprietary license.

    2026-09-25

    Release

  • Ember-1 is Fireworks Research's specialized foundation model announced September 23, 2026, the first model in the Ember series. Built on Kimi K3 and trained by Fireworks (50+ training experiments; Fireworks Serverless Training) to shorten reasoning traces (~40% fewer tokens while holding quality on Fireworks' evaluations and live A/B tests). Fireworks markets Ember-1 as its own model, a Research Preview on Serverless alongside base Kimi K3, and catalogs it as Fireworks Ember rather than a Kimi family member. Fireworks model page: API path accounts/fireworks/models/ember-1; MoE; parameters 2.78T; context length 1040k tokens; image input supported; function calling supported; fine-tuning not supported; serverless supported; kind base model; created 2026-09-22; state Ready; provider Fireworks. Serverless list is $3.00 input / $0.30 cached input / $15.00 output per 1M tokens. No first-party public weights or open license is stated.

    2026-09-23

    Release