ByteDance Doubao Seed 2.0 Pro
- τ-bench
- 90.4%
- Output (from)
- $2.37 / 1M
Last refreshed 2026-09-03. Next refresh: weekly.
Best LLMs for JSON / Tool use in 2026. Ranked by BFCL benchmark, with native JSON output and structured output support.
Verdict
Muse Spark 1.3 is the runner-up; compare τ-bench against Release.
JSON / Tool use leaders rank on BFCL first; when Berkeley coverage lags new SKUs, we fall back to τ-bench so fresh agentic models still surface.
| Model | Why it is listed | Status | Tracked price |
|---|---|---|---|
| Gemini 3.8 Flash ToolsCode execution | Gemini 3.8 Flash reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $0.75 / Out $3.75 |
| Muse Spark 1.3 Tools | Muse Spark 1.3 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $1.25 / Out $4.25 |
| Muse Spark 1.3 Contributor Tools | Muse Spark 1.3 Contributor reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $0.10 / Out $0.20 |
| Qwen3.8-Max-0902 Tools | Qwen3.8-Max-0902 reports tool-use capability; keep it separate from the scored tool-use ranking until benchmarks land. | Benchmark pending No tracked BFCL or tau-bench score yet. | In $2.00 / Out $6.00 |
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Ring-2.6-1T ReasoningTools Signal used: τ-bench 95.32% | $0.07 | $0.63 | |
| 2 | Mistral Medium 3.5 ReasoningVisionTools Signal used: τ-bench 91.4% | $1.50 | $7.50 | |
| 3 | ByteDance Doubao Seed 2.0 Pro VisionTools Signal used: τ-bench 90.4% | $0.47 | $2.37 | |
| 4 | Claude Mythos Preview Invite-onlyReasoningVisionTools Signal used: τ-bench 89.2% | — | — | |
| 5 | LFM2.5 8B A1B ReasoningTools Signal used: τ-bench 88.07% | — | — | |
| 6 | Claude Sonnet 4.6 ReasoningVisionTools Signal used: τ-bench 87.5% | $3.00 | $15.00 | |
| 7 | Command A+ ReasoningVisionTools Signal used: τ-bench 85% | — | — | |
| 8 | Claude Opus 4.6 ReasoningVisionTools Signal used: τ-bench 84.8% | $5.00 | $25.00 | |
| 9 | GLM-5 ReasoningTools Signal used: τ-bench 82.1% | $0.60 | $2.08 | |
| 10 | Qwen3.5-35B-A3B ReasoningTools Signal used: τ-bench 81.2% | $0.14 | $1.00 | |
| 11 | Qwen3.5-122B-A10B ReasoningVisionTools Signal used: τ-bench 79.5% | $0.26 | $2.08 | |
| 12 | Qwen3.5-9B VisionTools Signal used: τ-bench 79.1% | $0.10 | $0.15 | |
| 13 | Qwen3.5-27B ReasoningVisionTools Signal used: τ-bench 79% | $0.20 | $1.56 | |
| 14 | Grok 4.20 ReasoningVisionTools Signal used: τ-bench 78.9% | $1.25 | $2.50 | |
| 15 | GPT-5.4 ReasoningVisionTools Signal used: τ-bench 78.3% | $2.50 | $15.00 | |
| 16 | GPT-5.3-Codex ReasoningVisionTools Signal used: τ-bench 77.8% | $1.75 | $14.00 | |
| 17 | Claude Opus 4.5 ReasoningVisionTools Signal used: BFCL 77.47% | $5.00 | $25.00 | |
| 18 | Qwen3.6-Plus VisionTools Signal used: τ-bench 76.8% | $0.33 | $1.95 | |
| 19 | Qwen3-Max VisionTools Signal used: τ-bench 76.8% | $0.78 | $3.90 | |
| 20 | Gemini 3.1 Pro Preview PreviewVisionTools Signal used: τ-bench 76.5% | $2.00 | $12.00 |
Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8. Compare it for Coding, Agents, Long context, Vision, and JSON / Tool use.
2026-09-02
Release
Qwen3.8-Max-0902 is Alibaba's dated API snapshot of Qwen3.8-Max, listed September 2, 2026 on QwenCloud and Model Studio as model ID qwen3.8-max-0902 (first-party alias qwen3.8-max-2026-09-02). Official copy: an upgraded snapshot of qwen3.8-max with stronger coding (engineering-scale / long-horizon autonomous development), collaborative-agent multi-tool orchestration, and refined native vision (chart reasoning, document parsing, multimodal perception). Retains the 1,000,000-token context window, thinking mode, and full tool ecosystem. Input image/text/video, output text; function calling, structured outputs, context cache. International QwenCloud / Model Studio Singapore list price $2 input / $6 output per 1M tokens; implicit cache $0.25; explicit cache create $2.50; explicit cache read $0.17. No first-party 0902 weight drop. Sibling of seeded qwen3.8-max under family qwen3.8. Compare it for Agents and Coding.
2026-09-02
Release
Anthropic's generally available Mythos-class model for demanding reasoning, long-horizon agents, and ambitious coding, released September 1, 2026. Same underlying model as Claude Mythos 5.1 with cybersecurity and biology safeguards (flagged queries route to Opus 4.8 for cyber and Opus 5 for biology; those reroutes are not billed at Fable prices). Claude API id claude-fable-5-1. List price $10 per 1M input tokens and $50 per 1M output tokens; cache reads $0.25 per 1M (0.025x base input, 75% less than Fable 5's $1.00). 1M-token context, 128K max output, adaptive thinking always on (default effort high), text and image input to text output, vision, tool use. Available to Pro, Max, Team, and Enterprise users and on the Claude Platform natively, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Reliable knowledge cutoff June 2026. Sibling of seeded claude-fable-5 under family claude-fable. Compare it for Agents and Coding.
2026-09-01
Release