Kimi K3
- GPQA Diamond
- 93.5%
- Output (from)
- $11.20 / 1M
Last refreshed 2026-09-30. Next refresh: weekly.
The best open-weight LLMs in 2026, ranked by benchmark scores. Run locally, self-host, or deploy on your own infra — no API key required.
Verdict
Qwen3.8-Max is the runner-up, 0.9 points back on GPQA Diamond.
Open-weight boards emphasize GPQA Diamond (harder to game than broad MMLU), then MMLU, then recency.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Kimi K3 ReasoningVisionTools GPQA Diamond: 93.5% | $2.00 | $11.20 | |
| 2 | MiniMax M3 ReasoningVisionTools GPQA Diamond: 92.9% | $0.30 | $1.20 | |
| 3 | Qwen3.8-Max ReasoningVisionTools GPQA Diamond: 92.6% | $2.00 | $6.00 | |
| 4 | Qwen3.6-Max Vision GPQA Diamond: 91.8% | — | — | |
| 5 | Qwen3.8-Flash-Next ReasoningVision GPQA Diamond: 91.7% | $0.15 | $0.47 | |
| 6 | GLM-5.2 ReasoningTools GPQA Diamond: 91.2% | $1.40 | $4.40 | |
| 7 | DeepSeek V4.1 Flash ReasoningVisionTools GPQA Diamond: 90.9% | $0.15 | $0.60 | |
| 8 | Kimi K2.6 ReasoningVisionTools GPQA Diamond: 90.5% | $0.73 | $3.40 | |
| 9 | Qwen3.6-Plus VisionTools GPQA Diamond: 90.4% | $0.33 | $1.95 | |
| 10 | DeepSeek V4 Pro ReasoningTools GPQA Diamond: 90.1% | $0.43 | $0.87 | |
| 11 | Qwen3.5-397B-A17B ReasoningVisionTools GPQA Diamond: 89.3% | $0.39 | $2.34 | |
| 12 | Qwen3.8-27B ReasoningVisionTools GPQA Diamond: 89.2% | $0.21 | $2.55 | |
| 13 | Trinity-Large-Thinking ReasoningTools GPQA Diamond: 89.2% | $0.22 | $0.85 | |
| 14 | Qwen3.5-Plus Vision GPQA Diamond: 88.4% | $0.30 | $1.80 | |
| 15 | Ring-2.6-1T ReasoningTools GPQA Diamond: 88.27% | $0.07 | $0.63 | |
| 16 | DeepSeek V4 Flash ReasoningTools GPQA Diamond: 88.1% | $0.06 | $0.12 | |
| 17 | Qwen3.6-27B ReasoningVisionTools GPQA Diamond: 87.8% | $0.32 | $2.70 | |
| 18 | DeepSeek V3 0324 GPQA Diamond: 87.6% | $0.27 | $1.12 | |
| 19 | MiniMax M2.7 ReasoningTools GPQA Diamond: 87.4%Tied within margin | $0.28 | $1.20 | |
| 20 | Hunyuan Hy3 Preview PreviewReasoningTools GPQA Diamond: 87.2% | $0.07 | $0.26 |
GLM-5.2 is Z.ai's coding-first successor to GLM-5.1 in the GLM-5 family, released June 13 2026. 753B parameters (40B active) in IndexShare MoE architecture; the IndexShare innovation reuses the same attention indexer across every four sparse layers, cutting per-token FLOPs by 2.9x at 1M context length. Trained on 28.5T tokens. Supports a 1M-token context window via the glm-5.2[1m] model ID, with 131,072-token maximum output and High/Max thinking-effort levels designed for extended agentic coding sessions. MIT license; open weights available on Hugging Face (zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8). Self-reported HF card benchmarks: SWE-bench Pro 62.1, Terminal-Bench 2.1 82.7, MCP-Atlas 76.8, Tool-Decathlon 48.2, GPQA Diamond 91.2, AIME 2026 99.2, HLE 40.5. Available to GLM Coding Plan subscribers (Lite/Pro/Max/Team) directly, and via OpenRouter token API ($1.40/$4.40 per 1M tokens).
91.2%
GPQA Diamond
DeepSeek V4.1 Flash is a 552B-parameter (8B prefill / 16B decode activated) mixture-of-experts model with Compressed Expert Dispatch (CED), released September 10, 2026 under the MIT license. It supports 1M-token context, up to 384K output tokens, multimodal vision input, reasoning, function calling, tool use, structured outputs, and prompt caching. Weights are available on Hugging Face. The DeepSeek API serves it as deepseek-flash; legacy API aliases deepseek-v4-flash and deepseek-v4-flash-vision-exp route to this model. Off-peak API pricing: $0.15/1M input, $0.60/1M output (cache read: $0.003/1M); peak hours are 2×.
90.9%
GPQA Diamond
Kimi K2.6 is Moonshot AI's multimodal agentic coding model, released April 20 2026 under a Modified MIT license. Built on a 1-trillion-parameter MoE architecture (32B active, 384 experts with 8 selected per token plus 1 shared expert, 61 layers), it features a 262K context window and up to 65,536 output tokens. Supports native image and video inputs (screenshots, PDFs, spreadsheets). Designed for long-horizon coding with agent swarms of up to 300 sub-agents and 4,000 coordinated steps; Moonshot AI cites 200–300 sequential tool calls without task drift. Key benchmarks: SWE-bench Verified 80.2%, SWE-bench Pro 58.6%, LiveCodeBench v6 89.6%, GPQA Diamond 90.5%, Terminal-Bench 2.0 66.7%. Chatbot Arena Elo 1454 (2026-04-28 snapshot).
90.5%
GPQA Diamond