Best Reasoning LLMs (2026)

Last refreshed 2026-09-29. Next refresh: weekly.

Best reasoning and math LLMs in 2026, ranked by GPQA Diamond and MMLU. Includes thinking-mode models and chain-of-thought specialists.

Verdict

Use Kimi K3 for reasoning today.

Qwen3.8-Max is the runner-up, 0.9 points back on GPQA Diamond.

Researched 37d agoWhy this pickMethodology

How we rank

Reasoning boards prioritize GPQA Diamond scores, favoring models explicitly tagged for reasoning or unusually strong GPQA.

  1. Eligibility — Reasoning flag or GPQA Diamond above the editorial floor used on this page.
  2. Primary ranking — GPQA Diamond (higher is better), then newer release.
  3. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. Pricing — Reasoning tiers are often priced separately — confirm provider SKUs.
#ModelInput $/1MOutput $/1M
1GPT-5.6 Sol
ReasoningVisionTools

GPQA Diamond: 94.6%

$5.00$30.00
2Claude Mythos Preview
Invite-onlyReasoningVisionTools

GPQA Diamond: 94.6%

——
3Gemini 3.1 Pro Preview
PreviewVisionTools

GPQA Diamond: 94.3%

$2.00$12.00
4Claude Opus 4.7
ReasoningVisionTools

GPQA Diamond: 94.2%

$5.00$25.00
5Claude Opus 4.8
ReasoningVisionTools

GPQA Diamond: 93.6%

$5.00$25.00
6GPT-5.5
ReasoningVisionTools

GPQA Diamond: 93.6%

$5.00$30.00
7GPT-5.5 Pro
ReasoningVisionTools

GPQA Diamond: 93.6%

$30.00$180.00
8Kimi K3
ReasoningVisionTools

GPQA Diamond: 93.5%

$2.00$11.20
9MiniMax M3
ReasoningVisionTools

GPQA Diamond: 92.9%

$0.30$1.20
10Qwen3.8-Max
ReasoningVisionTools

GPQA Diamond: 92.6%

$2.00$6.00
11Qwen3.7-Max
ReasoningTools

GPQA Diamond: 92.4%

$1.25$3.75
12Gemini 3.5 Flash
ReasoningVisionTools

GPQA Diamond: 92.2%

$1.50$9.00
13GPT-5.4
ReasoningVisionTools

GPQA Diamond: 92%

$2.50$15.00
14Gemini 3 Pro
VisionTools

GPQA Diamond: 91.9%

$1.25$5.00
15Qwen3.6-Max
Vision

GPQA Diamond: 91.8%

——
16Qwen3.8-Flash-Next
ReasoningVision

GPQA Diamond: 91.7%

$0.15$0.47
17Claude Opus 4.6
ReasoningVisionTools

GPQA Diamond: 91.3%

$5.00$25.00
18GLM-5.2
ReasoningTools

GPQA Diamond: 91.2%

$1.40$4.40
19DeepSeek V4.1 Flash
ReasoningVisionTools

GPQA Diamond: 90.9%

$0.15$0.60
20Kimi K2.6
ReasoningVisionTools

GPQA Diamond: 90.5%

$0.73$3.40

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Qwen3.6-Max

    Qwen3.6-Max builds upon Qwen3-Max and Qwen3.6-Plus to deliver enhanced vibe coding capabilities, more efficient coding agent execution, and significantly improved front-end development skills. Its long-tail knowledge retention has been further upgraded, making it Alibaba's most capable multimodal flagship as of April 2026.

    91.8%

    GPQA Diamond

  • Qwen3.8-Flash-Next is Alibaba's experimental open-weight preview of the architecture planned for Qwen4. It is a 125B-total / 6B-active Mixture-of-Experts causal language model with a vision encoder, plus 51B n-gram embedding and 4B MTP parameters. Native context is 262,144 tokens (extensible to 1,000,000). Weights are on Hugging Face under the Qwen Community License 1.0. No first-party hosted token prices are seeded.

    91.7%

    GPQA Diamond

  • #6Claude Opus 4.6

    Claude Opus 4.6 is Anthropic's Claude 4.6 model with multimodal text and image input and an optional reasoning mode. It offers a 1M-token context window and scores 80.8 on SWE-bench Verified.

    91.3%

    GPQA Diamond