LLM Reference

Best LLMs for Code Generation (2026)

Last refreshed 2026-07-26. Next refresh: weekly.

The top coding LLMs in 2026, ranked by SWE-bench and HumanEval. Includes API pricing and context window for each pick — updated daily.

Verdict

Use Claude Fable 5 for code generation today.

Claude Opus 5 is the runner-up, 0.0 points back on SWE-bench Verified.

Researched 32d agoWhy this pickMethodology

How we rank

Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.

  1. EligibilityChat models tied to code work with current public/self-serve availability: code specialization, tracked code-execution flags, scores on SWE-bench / HumanEval / LiveCodeBench / Aider / BigCodeBench, or known code-family slugs.
  2. Primary rankingSWE-bench Verified (higher is better), then HumanEval, then SWE-bench Pro.
  3. Benchmark variantsIndependent standardized rows such as Vals.ai, Scale AI SWE-bench Pro, and Nebius/OpenHands SWE-bench Verified stay separate from vendor-reported variants. Vendor-only orchestration claims do not promote a model on this page until a comparable public harness publishes the score.
  4. Tie-breaksNewer `release` date when benchmark scores match.
  5. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  6. Preview fallbackThe podium prefers GA models. Preview or invite-only candidates can fill this page only when fewer than three GA coding primaries remain after the gate.
  7. Pricing columnInput/output prices are the lowest tracked public commercial rate cards in seed data; partner-only pricing is kept out of the default ranking.

New models awaiting benchmark coverage

These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.

ModelWhy it is listedStatusTracked price
Gemini 3.5 Flash-Lite
ToolsCode execution
Gemini 3.5 Flash-Lite is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $0.30 / Out $2.50
Gemini 3.6 Flash
ToolsCode execution
Gemini 3.6 Flash is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $1.50 / Out $7.50
Kimi K2.7-Code HighSpeed
Tools
Kimi K2.7-Code HighSpeed is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $1.90 / Out $8.00
Kimi K2.7-Code
Tools
Kimi K2.7-Code is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $0.61 / Out $3.07
#ModelInput $/1MOutput $/1M
1Claude Fable 5
ReasoningVisionTools

SWE-bench Verified: 96%

$10.00$50.00
2Claude Opus 5
ReasoningVisionTools

SWE-bench Verified: 96%

$5.00$25.00
3Claude Opus 4.8
ReasoningVisionTools

SWE-bench Verified: 88.6%

$5.00$25.00
4Claude Opus 4.7
ReasoningVisionTools

SWE-bench Verified: 87.6%

$5.00$25.00
5Claude Sonnet 5
ReasoningVisionTools

SWE-bench Verified: 85.2%

$2.00$10.00
6GPT-5.3-Codex
ReasoningVisionTools

SWE-bench Verified: 85%

$1.75$14.00
7GPT-5.5
ReasoningVisionTools

SWE-bench Verified: 82.6%

$5.00$30.00
8GPT-5.5 Pro
ReasoningVisionTools

SWE-bench Verified: 82.6%

$30.00$180.00
9Claude Opus 4.5
ReasoningVisionTools

SWE-bench Verified: 80.9%

$5.00$25.00
10Claude Opus 4.6
ReasoningVisionTools

SWE-bench Verified: 80.8%

$5.00$25.00
11Gemini 3.1 Pro Preview
PreviewVisionTools

SWE-bench Verified: 80.6%

$2.00$12.00
12DeepSeek V4 Pro
ReasoningTools

SWE-bench Verified: 80.6%

$0.43$0.87
13MiniMax M3
ReasoningVisionTools

SWE-bench Verified: 80.5%

$0.30$1.20
14Qwen3.7-Max
ReasoningTools

SWE-bench Verified: 80.4%

$1.25$3.75
15Kimi K2.6
ReasoningVisionTools

SWE-bench Verified: 80.2%

$0.73$3.40
16MiniMax M2.5 Highspeed
ReasoningTools

SWE-bench Verified: 80.2%

$0.60$2.40
17Claude Sonnet 4.6
ReasoningVisionTools

SWE-bench Verified: 79.6%

$3.00$15.00
18DeepSeek V4 Flash
ReasoningTools

SWE-bench Verified: 79%

$0.09$0.18
19Xiaomi MiMo-V2.5-Pro
Tools

SWE-bench Verified: 78.9%

$0.43$0.87
20Qwen3.6 Max Preview
PreviewReasoningVisionTools

SWE-bench Verified: 78.8%

$1.04$6.24

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.

  • Claude Opus 4.7 is Anthropic's generally available flagship model with 1M context, 128K max output, adaptive thinking, and a new tokenizer with roughly 555K words per 1M tokens.

    87.6%

    SWE-bench Verified

  • Claude Sonnet 5 is Anthropic's next-generation Sonnet model for agentic coding, tool use, computer use, and professional work. It is a proprietary decoder-only model with a 1M-token context window, 128K max output, multimodal vision, adaptive thinking, function calling, structured outputs, prompt caching, and Batch API support. It is available through the Claude API, AWS Bedrock, Google Cloud Vertex AI, Microsoft Foundry preview, and OpenRouter. Anthropic lists durable standard pricing at $3/1M input and $15/1M output tokens, with introductory $2/$10 pricing through 2026-08-31.

    85.2%

    SWE-bench Verified

  • #6GPT-5.3-Codex

    Most capable agentic coding model from OpenAI. Optimized for long-horizon, agentic coding tasks in the Codex CLI and API. Note: GPT-5.3-Codex-Spark is a distinct ChatGPT Pro research preview (not API-accessible).

    85%

    SWE-bench Verified

Frequently asked questions

Which LLM is best for code generation?

Claude Fable 5 is the current LLMReference top pick for code generation. The verdict uses the stored category signal SWE-bench Verified: 96%. Output pricing starts at $50.00 per 1M tokens. Review the linked model and provider pages before production use because availability and pricing can change.

How does Claude Fable 5 compare to Claude Opus 5 for code generation?

Claude Fable 5 leads Claude Opus 5 in the visible shortlist on SWE-bench Verified: 96% versus 96%. The pricing cards show Claude Fable 5: output pricing starts at $50.00 per 1m tokens and Claude Opus 5: output pricing starts at $25.00 per 1m tokens.

How does LLMReference rank LLMs for code generation?

LLMReference ranks LLMs for code generation from stored model, benchmark, freshness, and pricing data. The current methodology summary is: Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.

How often is this list updated?

The LLM rankings on this page are updated daily as new benchmark scores, provider availability, and pricing data are tracked. The "as of" date at the top of the page shows the most recent refresh.

How do you decide which models appear in the top 3?

The podium picks are driven by the primary benchmark signal for this category (shown in the Methodology section), filtered to non-deprecated models with confirmed API availability. In ties, we prefer the more recently released model.

Are preview or beta models included?

Preview models appear in the "Watch list" section but are not in the main ranked podium unless the category explicitly allows it (e.g., /best/coding and /best/agents, where preview models often lead benchmarks).

Can I compare two specific models head-to-head?

Yes — use the Compare tool at llmreference.com/compare for a side-by-side breakdown of context window, pricing, benchmarks, and provider availability.

Is the pricing data real-time?

Pricing is tracked from provider documentation and updated regularly. It reflects the best available public data, not live API quotes — always verify before billing.