Kimi K3
- GPQA Diamond
- 93.5%
- Output (from)
- $15.00 / 1M
Last refreshed 2026-09-01. Next refresh: weekly.
The best open-weight LLMs in 2026, ranked by benchmark scores. Run locally, self-host, or deploy on your own infra — no API key required.
Verdict
Qwen3.8-Max is the runner-up, 0.9 points back on GPQA Diamond.
Open-weight boards emphasize GPQA Diamond (harder to game than broad MMLU), then MMLU, then recency.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Kimi K3 ReasoningVisionTools GPQA Diamond: 93.5% | $3.00 | $15.00 | |
| 2 | MiniMax M3 ReasoningVisionTools GPQA Diamond: 92.9% | $0.30 | $1.20 | |
| 3 | Qwen3.8-Max ReasoningVisionTools GPQA Diamond: 92.6% | $2.00 | $6.00 | |
| 4 | Qwen3.6-Max Vision GPQA Diamond: 91.8% | — | — | |
| 5 | Qwen3.8-Flash-Next ReasoningVision GPQA Diamond: 91.7% | — | — | |
| 6 | GLM-5.2 ReasoningTools GPQA Diamond: 91.2% | $1.40 | $4.40 | |
| 7 | Kimi K2.6 ReasoningVisionTools GPQA Diamond: 90.5% | $0.73 | $3.40 | |
| 8 | Qwen3.6-Plus VisionTools GPQA Diamond: 90.4% | $0.33 | $1.95 | |
| 9 | DeepSeek V4 Pro ReasoningTools GPQA Diamond: 90.1% | $0.43 | $0.87 | |
| 10 | Qwen3.5-397B-A17B ReasoningVisionTools GPQA Diamond: 89.3% | $0.39 | $2.34 | |
| 11 | Qwen3.8-27B ReasoningVisionTools GPQA Diamond: 89.2% | $0.50 | $3.00 | |
| 12 | Trinity-Large-Thinking ReasoningTools GPQA Diamond: 89.2% | $0.22 | $0.85 | |
| 13 | Qwen3.5-Plus Vision GPQA Diamond: 88.4% | $0.30 | $1.80 | |
| 14 | Ring-2.6-1T ReasoningTools GPQA Diamond: 88.27% | $0.07 | $0.63 | |
| 15 | DeepSeek V4 Flash ReasoningTools GPQA Diamond: 88.1% | $0.06 | $0.12 | |
| 16 | Qwen3.6-27B ReasoningVisionTools GPQA Diamond: 87.8% | $0.32 | $3.20 | |
| 17 | DeepSeek V3 0324 GPQA Diamond: 87.6% | $0.27 | $1.12 | |
| 18 | MiniMax M2.7 ReasoningTools GPQA Diamond: 87.4%Tied within margin | $0.28 | $1.20 | |
| 19 | Hunyuan Hy3 Preview PreviewReasoningTools GPQA Diamond: 87.2% | $0.07 | $0.26 | |
| 20 | GLM-5.1 ReasoningTools GPQA Diamond: 86.2% | $1.05 | $3.50 |
Arcee AI's flagship 400B sparse MoE reasoning model with 13B active parameters per token. Trained on 20T tokens with a STEM-focused curriculum. Designed for agentic workflows, chain-of-thought reasoning, and long-context tasks up to 256K tokens (BF16 API). Open-source under Apache 2.0. Available via Arcee AI API.
89.2%
GPQA Diamond
Qwen3.5-Plus is the flagship commercial API model of the Qwen3.5 native vision-language series, delivering outstanding performance comparable to state-of-the-art models with significant leaps in both pure-text and multimodal capabilities compared to the Qwen3 series.
88.4%
GPQA Diamond
Ring-2.6-1T is InclusionAI's MIT-licensed trillion-parameter MoE reasoning model for agent workflows, engineering tasks, scientific analysis, and enterprise automation. It supports high and xhigh reasoning effort modes and entered OpenRouter's Programming top 10 in the 2026-05-18 audit.
88.27%
GPQA Diamond