Gemini 3 Pro
- SWE-bench Verified
- 76.2%
- Output (from)
- $5.00 / 1M
Last refreshed 2026-09-29. Next refresh: weekly.
The top coding LLMs in 2026, ranked by SWE-bench and HumanEval. Includes API pricing and context window for each pick — updated daily.
Verdict
Phi-3 Mini 128K is the runner-up: 76.2% vs — on SWE-bench Verified.
Coding leaders are ordered on shipped coding-agent evidence first, then classic code generation scores, with recency as the last tie-break.
| Model | Why it is listed | Status | Tracked price |
|---|---|---|---|
| GPT-6.1 Sol ToolsCode execution | GPT-6.1 Sol is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $2.00 / Out $10.00 |
| GPT-6 Luna ToolsCode execution | GPT-6 Luna is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $0.10 / Out $0.50 |
| GPT-6 Sol ToolsCode execution | GPT-6 Sol is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $2.00 / Out $10.00 |
| Grok 4.7 ToolsCode execution | Grok 4.7 is a newly researched coding-capable model; keep it on the watchlist until category scores land. | Benchmark pending No tracked SWE-bench Verified, HumanEval, or SWE-bench Pro score yet. | In $1.60 / Out $4.80 |
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Claude Fable 5 ReasoningVisionTools SWE-bench Verified: 96% | $10.00 | $50.00 | |
| 2 | Claude Opus 5 ReasoningVisionTools SWE-bench Verified: 96% | $5.00 | $25.00 | |
| 3 | Claude Opus 4.8 ReasoningVisionTools SWE-bench Verified: 88.6% | $5.00 | $25.00 | |
| 4 | Claude Opus 4.7 ReasoningVisionTools SWE-bench Verified: 87.6% | $5.00 | $25.00 | |
| 5 | Claude Sonnet 5 ReasoningVisionTools SWE-bench Verified: 85.2% | $2.00 | $10.00 | |
| 6 | GPT-5.3-Codex ReasoningVisionTools SWE-bench Verified: 85% | $1.75 | $14.00 | |
| 7 | GPT-5.5 ReasoningVisionTools SWE-bench Verified: 82.6% | $5.00 | $30.00 | |
| 8 | GPT-5.5 Pro ReasoningVisionTools SWE-bench Verified: 82.6% | $30.00 | $180.00 | |
| 9 | Claude Opus 4.5 ReasoningVisionTools SWE-bench Verified: 80.9% | $5.00 | $25.00 | |
| 10 | Claude Opus 4.6 ReasoningVisionTools SWE-bench Verified: 80.8% | $5.00 | $25.00 | |
| 11 | Gemini 3.1 Pro Preview PreviewVisionTools SWE-bench Verified: 80.6% | $2.00 | $12.00 | |
| 12 | DeepSeek V4 Pro ReasoningTools SWE-bench Verified: 80.6% | $0.43 | $0.87 | |
| 13 | MiniMax M3 ReasoningVisionTools SWE-bench Verified: 80.5% | $0.30 | $1.20 | |
| 14 | Qwen3.7-Max ReasoningTools SWE-bench Verified: 80.4% | $1.25 | $3.75 | |
| 15 | Kimi K2.6 ReasoningVisionTools SWE-bench Verified: 80.2% | $0.73 | $3.40 | |
| 16 | MiniMax M2.5 Highspeed ReasoningTools SWE-bench Verified: 80.2% | $0.60 | $2.40 | |
| 17 | Claude Sonnet 4.6 ReasoningVisionTools SWE-bench Verified: 79.6% | $3.00 | $15.00 | |
| 18 | DeepSeek V4 Flash ReasoningTools SWE-bench Verified: 79% | $0.06 | $0.12 | |
| 19 | Xiaomi MiMo-V2.5-Pro Tools SWE-bench Verified: 78.9% | $0.43 | $0.87 | |
| 20 | Qwen3.6 Max Preview PreviewReasoningVisionTools SWE-bench Verified: 78.8% | $1.04 | $6.24 |
Aya-23-8B is a multilingual large language model developed by Cohere For AI, featuring 8 billion parameters. As an instruction-fine-tuned model, it is adept at following instructions and is optimized for text generation and understanding. The model employs a decoder-only Transformer architecture, utilizing enhancements like parallel attention and feed-forward layers for efficiency. It supports 23 languages, including Arabic, Chinese, English, and French, and is proficient in tasks such as machine translation, chatbot interactions, and text summarization. Despite its capabilities, its performance might vary across languages, particularly those with less linguistic resources, and it has a context length limit of 8192 tokens. Training involved diverse data sources like human annotations and synthetic datasets to bolster its multilingual proficiency.
—
SWE-bench Verified
Instruction-tuned 7B variant combining strong reasoning with real-time inference on single GPUs, ideal for developer tools and vision applications.
—
SWE-bench Verified
The Llama 3 8B Instruct model, released on April 18, 2024, is Meta's latest instruction-following language model with 8 billion parameters. It utilizes an auto-regressive transformer architecture with Grouped-Query Attention for improved scalability. Trained on over 15 trillion tokens and fine-tuned with 10 million human-annotated examples, it excels in dialogue and conversational tasks. The model outperforms its predecessors on industry benchmarks, scoring 68.4 on MMLU (5-shot). Designed for commercial and research applications, it prioritizes safety and helpfulness, making it suitable for chatbots, virtual assistants, and other interactive AI applications. For more details, visit the Hugging Face page [1].
—
SWE-bench Verified