Claude Fable 5
- SWE-bench Verified
- 96%
- Output (from)
- $50.00 / 1M
Last refreshed 2026-07-26. Next refresh: weekly.
The top coding LLMs in 2026, ranked by SWE-bench and HumanEval. Includes API pricing and context window for each pick — updated daily.
Verdict
Claude Opus 5 is the runner-up, 0.0 points back on SWE-bench Verified.
Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.
These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.
| Model | Why it is listed | Status | Tracked price |
|---|---|---|---|
| Gemini 3.5 Flash-Lite ToolsCode execution | Gemini 3.5 Flash-Lite is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $0.30 / Out $2.50 |
| Gemini 3.6 Flash ToolsCode execution | Gemini 3.6 Flash is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $1.50 / Out $7.50 |
| Kimi K2.7-Code HighSpeed Tools | Kimi K2.7-Code HighSpeed is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $1.90 / Out $8.00 |
| Kimi K2.7-Code Tools | Kimi K2.7-Code is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $0.61 / Out $3.07 |
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Claude Fable 5 ReasoningVisionTools SWE-bench Verified: 96% | $10.00 | $50.00 | |
| 2 | Claude Opus 5 ReasoningVisionTools SWE-bench Verified: 96% | $5.00 | $25.00 | |
| 3 | Claude Opus 4.8 ReasoningVisionTools SWE-bench Verified: 88.6% | $5.00 | $25.00 | |
| 4 | Claude Opus 4.7 ReasoningVisionTools SWE-bench Verified: 87.6% | $5.00 | $25.00 | |
| 5 | Claude Sonnet 5 ReasoningVisionTools SWE-bench Verified: 85.2% | $2.00 | $10.00 | |
| 6 | GPT-5.3-Codex ReasoningVisionTools SWE-bench Verified: 85% | $1.75 | $14.00 | |
| 7 | GPT-5.5 ReasoningVisionTools SWE-bench Verified: 82.6% | $5.00 | $30.00 | |
| 8 | GPT-5.5 Pro ReasoningVisionTools SWE-bench Verified: 82.6% | $30.00 | $180.00 | |
| 9 | Claude Opus 4.5 ReasoningVisionTools SWE-bench Verified: 80.9% | $5.00 | $25.00 | |
| 10 | Claude Opus 4.6 ReasoningVisionTools SWE-bench Verified: 80.8% | $5.00 | $25.00 | |
| 11 | Gemini 3.1 Pro Preview PreviewVisionTools SWE-bench Verified: 80.6% | $2.00 | $12.00 | |
| 12 | DeepSeek V4 Pro ReasoningTools SWE-bench Verified: 80.6% | $0.43 | $0.87 | |
| 13 | MiniMax M3 ReasoningVisionTools SWE-bench Verified: 80.5% | $0.30 | $1.20 | |
| 14 | Qwen3.7-Max ReasoningTools SWE-bench Verified: 80.4% | $1.25 | $3.75 | |
| 15 | Kimi K2.6 ReasoningVisionTools SWE-bench Verified: 80.2% | $0.73 | $3.40 | |
| 16 | MiniMax M2.5 Highspeed ReasoningTools SWE-bench Verified: 80.2% | $0.60 | $2.40 | |
| 17 | Claude Sonnet 4.6 ReasoningVisionTools SWE-bench Verified: 79.6% | $3.00 | $15.00 | |
| 18 | DeepSeek V4 Flash ReasoningTools SWE-bench Verified: 79% | $0.09 | $0.18 | |
| 19 | Xiaomi MiMo-V2.5-Pro Tools SWE-bench Verified: 78.9% | $0.43 | $0.87 | |
| 20 | Qwen3.6 Max Preview PreviewReasoningVisionTools SWE-bench Verified: 78.8% | $1.04 | $6.24 |
Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
Claude Opus 4.7 is Anthropic's generally available flagship model with 1M context, 128K max output, adaptive thinking, and a new tokenizer with roughly 555K words per 1M tokens.
87.6%
SWE-bench Verified
Claude Sonnet 5 is Anthropic's next-generation Sonnet model for agentic coding, tool use, computer use, and professional work. It is a proprietary decoder-only model with a 1M-token context window, 128K max output, multimodal vision, adaptive thinking, function calling, structured outputs, prompt caching, and Batch API support. It is available through the Claude API, AWS Bedrock, Google Cloud Vertex AI, Microsoft Foundry preview, and OpenRouter. Anthropic lists durable standard pricing at $3/1M input and $15/1M output tokens, with introductory $2/$10 pricing through 2026-08-31.
85.2%
SWE-bench Verified
Most capable agentic coding model from OpenAI. Optimized for long-horizon, agentic coding tasks in the Codex CLI and API. Note: GPT-5.3-Codex-Spark is a distinct ChatGPT Pro research preview (not API-accessible).
85%
SWE-bench Verified
Side-by-side comparison of the top picks by price, benchmark, and API access.
Claude Fable 5 is the current LLMReference top pick for code generation. The verdict uses the stored category signal SWE-bench Verified: 96%. Output pricing starts at $50.00 per 1M tokens. Review the linked model and provider pages before production use because availability and pricing can change.
Claude Fable 5 leads Claude Opus 5 in the visible shortlist on SWE-bench Verified: 96% versus 96%. The pricing cards show Claude Fable 5: output pricing starts at $50.00 per 1m tokens and Claude Opus 5: output pricing starts at $25.00 per 1m tokens.
LLMReference ranks LLMs for code generation from stored model, benchmark, freshness, and pricing data. The current methodology summary is: Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.
The LLM rankings on this page are updated daily as new benchmark scores, provider availability, and pricing data are tracked. The "as of" date at the top of the page shows the most recent refresh.
The podium picks are driven by the primary benchmark signal for this category (shown in the Methodology section), filtered to non-deprecated models with confirmed API availability. In ties, we prefer the more recently released model.
Preview models appear in the "Watch list" section but are not in the main ranked podium unless the category explicitly allows it (e.g., /best/coding and /best/agents, where preview models often lead benchmarks).
Yes — use the Compare tool at llmreference.com/compare for a side-by-side breakdown of context window, pricing, benchmarks, and provider availability.
Pricing is tracked from provider documentation and updated regularly. It reflects the best available public data, not live API quotes — always verify before billing.