Gemini 3 Pro
- BFCL
- 72.51%
- Output (from)
- $5.00 / 1M
Last refreshed 2026-09-29. Next refresh: weekly.
Best LLMs for JSON / Tool use in 2026. Ranked by BFCL benchmark, with native JSON output and structured output support.
Verdict
GPT-6.1 Sol is the runner-up; compare BFCL against Release.
JSON / Tool use leaders rank on BFCL first; when Berkeley coverage lags new SKUs, we fall back to τ-bench so fresh agentic models still surface.
| Model | Why it is listed | Status | Tracked price |
|---|---|---|---|
| GPT-6.1 Sol ToolsCode execution | GPT-6.1 Sol reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $2.00 / Out $10.00 |
| Claude Sonnet 5.5 Tools | Claude Sonnet 5.5 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $2.00 / Out $10.00 |
| Perceptron Mk1.5 Tools | Perceptron Mk1.5 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $0.15 / Out $1.50 |
| Ember-1 Tools | Ember-1 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $3.00 / Out $15.00 |
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Ring-2.6-1T ReasoningTools Signal used: τ-bench 95.32% | $0.07 | $0.63 | |
| 2 | Mistral Medium 3.5 ReasoningVisionTools Signal used: τ-bench 91.4% | $1.50 | $7.50 | |
| 3 | ByteDance Doubao Seed 2.0 Pro VisionTools Signal used: τ-bench 90.4% | $0.47 | $2.37 | |
| 4 | Claude Mythos Preview Invite-onlyReasoningVisionTools Signal used: τ-bench 89.2% | — | — | |
| 5 | LFM2.5 8B A1B ReasoningTools Signal used: τ-bench 88.07% | — | — | |
| 6 | Claude Sonnet 4.6 ReasoningVisionTools Signal used: τ-bench 87.5% | $3.00 | $15.00 | |
| 7 | Command A+ ReasoningVisionTools Signal used: τ-bench 85% | — | — | |
| 8 | Claude Opus 4.6 ReasoningVisionTools Signal used: τ-bench 84.8% | $5.00 | $25.00 | |
| 9 | GLM-5 ReasoningTools Signal used: τ-bench 82.1% | $0.60 | $2.08 | |
| 10 | Qwen3.5-35B-A3B ReasoningTools Signal used: τ-bench 81.2% | $0.14 | $1.00 | |
| 11 | Qwen3.5-122B-A10B ReasoningVisionTools Signal used: τ-bench 79.5% | $0.26 | $2.08 | |
| 12 | Qwen3.5-9B VisionTools Signal used: τ-bench 79.1% | $0.10 | $0.15 | |
| 13 | Qwen3.5-27B ReasoningVisionTools Signal used: τ-bench 79% | $0.20 | $1.56 | |
| 14 | Grok 4.20 ReasoningVisionTools Signal used: τ-bench 78.9% | $1.25 | $2.50 | |
| 15 | GPT-5.4 ReasoningVisionTools Signal used: τ-bench 78.3% | $2.50 | $15.00 | |
| 16 | GPT-5.3-Codex ReasoningVisionTools Signal used: τ-bench 77.8% | $1.75 | $14.00 | |
| 17 | Claude Opus 4.5 ReasoningVisionTools Signal used: BFCL 77.47% | $5.00 | $25.00 | |
| 18 | Qwen3.6-Plus VisionTools Signal used: τ-bench 76.8% | $0.33 | $1.95 | |
| 19 | Qwen3-Max VisionTools Signal used: τ-bench 76.8% | $0.78 | $3.90 | |
| 20 | Gemini 3.1 Pro Preview PreviewVisionTools Signal used: τ-bench 76.5% | $2.00 | $12.00 |
MiniMax M3.1 Flash Preview (API id MiniMax-M3.1-Flash-Preview) is MiniMax's faster/lighter M3-series multimodal coding model announced for MiniMax Code on 2026-09-27 and made live on Token Plan the same day. First-party docs: 1,000,000-token context; text+image+video inputs to text; agentic reasoning, tool use, and coding; thinking always on with tunable effort low|medium|high|xhigh|max (default max; cannot disable — 400 'requires adaptive thinking'). Availability for now is Token Plan + MiniMax Code only — NOT on pay-as-you-go (paygo price table lists M3/M2.x only; no M3.1 Flash $/MTok row). No public Hugging Face weights/model card for the Flash Preview SKU. Closed-source hosted API — licenseSlug proprietary / weights notReleased.
2026-09-27
Release
Perceptron Mk1.5 is Perceptron's embodied vision-language model released September 25, 2026 (first-party blog). Successor to Mk1 in the Perceptron Mk family: native audio in, video object tracking, web search, sub-agent/tool calls, and stronger visual reasoning; blog cites ~2-5x faster end-to-end vs Mk1. Accepts text, images, video, and audio; emits text plus spatial/temporal annotations (points, boxes, polygons, clips, object tracks). First-party docs: Model ID perceptron-mk1.5; context window 36,864 tokens; max output 8,192; pricing $0.15/M input, $1.50/M output, $0.0375/M cached input. Available via Perceptron Platform + SDK and OpenRouter (perceptron/perceptron-mk1.5). Closed-source hosted API under a proprietary license.
2026-09-25
Release
Ember-1 is Fireworks Research's specialized foundation model announced September 23, 2026, the first model in the Ember series. Built on Kimi K3 and trained by Fireworks (50+ training experiments; Fireworks Serverless Training) to shorten reasoning traces (~40% fewer tokens while holding quality on Fireworks' evaluations and live A/B tests). Fireworks markets Ember-1 as its own model, a Research Preview on Serverless alongside base Kimi K3, and catalogs it as Fireworks Ember rather than a Kimi family member. Fireworks model page: API path accounts/fireworks/models/ember-1; MoE; parameters 2.78T; context length 1040k tokens; image input supported; function calling supported; fine-tuning not supported; serverless supported; kind base model; created 2026-09-22; state Ready; provider Fireworks. Serverless list is $3.00 input / $0.30 cached input / $15.00 output per 1M tokens. No first-party public weights or open license is stated.
2026-09-23
Release