Best LLMs for Code Generation (2026)

Last refreshed 2026-09-29. Next refresh: weekly.

The top coding LLMs in 2026, ranked by SWE-bench and HumanEval. Includes API pricing and context window for each pick — updated daily.

Verdict

Use Gemini 3 Pro for code generation today.

Phi-3 Mini 128K is the runner-up: 76.2% vs — on SWE-bench Verified.

Researched 8d agoWhy this pickMethodology

How we rank

Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.

  1. Eligibility — Chat models tied to code work with current public/self-serve availability: code specialization, tracked code-execution flags, scores on SWE-bench / HumanEval / LiveCodeBench / Aider / BigCodeBench, or known code-family slugs.
  2. Primary ranking — SWE-bench Verified (higher is better), then HumanEval, then SWE-bench Pro.
  3. Benchmark variants — Independent standardized rows such as Vals.ai, Scale AI SWE-bench Pro, and Nebius/OpenHands SWE-bench Verified stay separate from vendor-reported variants. Vendor-only orchestration claims do not promote a model on this page until a comparable public harness publishes the score.
  4. Tie-breaks — Newer `release` date when benchmark scores match.
  5. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  6. Preview fallback — The podium prefers GA models. Preview or invite-only candidates can fill this page only when fewer than three GA coding primaries remain after the gate.
  7. Pricing column — Input/output prices are the lowest tracked public commercial rate cards in seed data; partner-only pricing is kept out of the default ranking.

New models awaiting benchmark coverage

These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.
ModelWhy it is listedStatusTracked price
GPT-6.1 Sol
ToolsCode execution
GPT-6.1 Sol is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $2.00 / Out $10.00
GPT-6 Luna
ToolsCode execution
GPT-6 Luna is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $0.10 / Out $0.50
GPT-6 Sol
ToolsCode execution
GPT-6 Sol is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $2.00 / Out $10.00
Grok 4.7
ToolsCode execution
Grok 4.7 is a newly researched coding-capable model; keep it on the watchlist until category scores land.Benchmark pending

No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet.

In $1.60 / Out $4.80
#ModelInput $/1MOutput $/1M
1Claude Fable 5
ReasoningVisionTools

SWE-bench Verified: 96%

$10.00$50.00
2Claude Opus 5
ReasoningVisionTools

SWE-bench Verified: 96%

$5.00$25.00
3Claude Opus 4.8
ReasoningVisionTools

SWE-bench Verified: 88.6%

$5.00$25.00
4Claude Opus 4.7
ReasoningVisionTools

SWE-bench Verified: 87.6%

$5.00$25.00
5Claude Sonnet 5
ReasoningVisionTools

SWE-bench Verified: 85.2%

$2.00$10.00
6GPT-5.3-Codex
ReasoningVisionTools

SWE-bench Verified: 85%

$1.75$14.00
7GPT-5.5
ReasoningVisionTools

SWE-bench Verified: 82.6%

$5.00$30.00
8GPT-5.5 Pro
ReasoningVisionTools

SWE-bench Verified: 82.6%

$30.00$180.00
9Claude Opus 4.5
ReasoningVisionTools

SWE-bench Verified: 80.9%

$5.00$25.00
10Claude Opus 4.6
ReasoningVisionTools

SWE-bench Verified: 80.8%

$5.00$25.00
11Gemini 3.1 Pro Preview
PreviewVisionTools

SWE-bench Verified: 80.6%

$2.00$12.00
12DeepSeek V4 Pro
ReasoningTools

SWE-bench Verified: 80.6%

$0.43$0.87
13MiniMax M3
ReasoningVisionTools

SWE-bench Verified: 80.5%

$0.30$1.20
14Qwen3.7-Max
ReasoningTools

SWE-bench Verified: 80.4%

$1.25$3.75
15Kimi K2.6
ReasoningVisionTools

SWE-bench Verified: 80.2%

$0.73$3.40
16MiniMax M2.5 Highspeed
ReasoningTools

SWE-bench Verified: 80.2%

$0.60$2.40
17Claude Sonnet 4.6
ReasoningVisionTools

SWE-bench Verified: 79.6%

$3.00$15.00
18DeepSeek V4 Flash
ReasoningTools

SWE-bench Verified: 79%

$0.06$0.12
19Xiaomi MiMo-V2.5-Pro
Tools

SWE-bench Verified: 78.9%

$0.43$0.87
20Qwen3.6 Max Preview
PreviewReasoningVisionTools

SWE-bench Verified: 78.8%

$1.04$6.24

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Aya-23-8B is a multilingual large language model developed by Cohere For AI, featuring 8 billion parameters. As an instruction-fine-tuned model, it is adept at following instructions and is optimized for text generation and understanding. The model employs a decoder-only Transformer architecture, utilizing enhancements like parallel attention and feed-forward layers for efficiency. It supports 23 languages, including Arabic, Chinese, English, and French, and is proficient in tasks such as machine translation, chatbot interactions, and text summarization. Despite its capabilities, its performance might vary across languages, particularly those with less linguistic resources, and it has a context length limit of 8192 tokens. Training involved diverse data sources like human annotations and synthetic datasets to bolster its multilingual proficiency.

    —

    SWE-bench Verified

  • #5Qwen2.5-7B-Instruct

    Instruction-tuned 7B variant combining strong reasoning with real-time inference on single GPUs, ideal for developer tools and vision applications.

    —

    SWE-bench Verified

  • #6Llama 3 8B Instruct

    The Llama 3 8B Instruct model, released on April 18, 2024, is Meta's latest instruction-following language model with 8 billion parameters. It utilizes an auto-regressive transformer architecture with Grouped-Query Attention for improved scalability. Trained on over 15 trillion tokens and fine-tuned with 10 million human-annotated examples, it excels in dialogue and conversational tasks. The model outperforms its predecessors on industry benchmarks, scoring 68.4 on MMLU (5-shot). Designed for commercial and research applications, it prioritizes safety and helpfulness, making it suitable for chatbots, virtual assistants, and other interactive AI applications. For more details, visit the Hugging Face page [1].

    —

    SWE-bench Verified