LLM Reference

Compare AI models

Side-by-side comparison of any two LLMs — GPT vs Claude, Gemini vs DeepSeek, open vs proprietary — on pricing, benchmarks, API availability, context window, and release date.

Sitemap coverage 4408+ pairs

Decision builder

Pick the pair before opening the detail page

220 selectable models
Open comparison

Claude Opus 4.7 vs Claude Opus 4.8

Pick Claude Opus 4.8 for higher current agentic coding and computer-use confidence; token pricing is tied on tracked $5/1M input and $25/1M output routes, so keep Claude Opus 4.7 only for already-validated prompts or coding workflow support constraints.

0% gap
Output price
$25.00 / $25.00
Context
1m / 1m
Benchmarks
5 shared
Providers
6 / 6

Popular pairs

Browse comparisons with a decision signal attached

GPT-5.6 Sol vs Grok 4.5

Pick GPT-5.6 Sol when you want OpenAI's July 2026 frontier stack, the larger 1.05M context window, and OpenAI-sourced GA rows such as DeepSWE 1.1 at 72.7% and GPQA Diamond at 94.6%. Pick Grok 4.5 when xAI's lower standard-tier API pricing ($2/$6 per 1M tokens for prompts up to 200K) and Grok Build/Cursor distribution matter more, and run your own acceptance tests because several Grok 4.5 chart scores are xAI first-party only. Do not treat Terra or Luna as the OpenAI flagship in this pair.

400% gap8 benchmarks
Output price
$30.00 / $6.00
Context
1.05m / 500k
Benchmarks
8 shared
Providers
2 / 3
CodingRAGAgentsLong contextGrok 4.5 leads SWE-bench Pro

Claude Fable 5 vs GPT-5.6 Sol

Pick GPT-5.6 Sol for a fully available OpenAI GA route with 1.05M context, lower standard output pricing at $30/M versus Fable 5's $50/M, and OpenAI launch rows you can cite directly in procurement reviews. Pick Claude Fable 5 when Anthropic's agentic coding evidence, adaptive thinking, and long-horizon workflow positioning outweigh access verification, especially after you confirm live Fable 5 availability on your provider route. Do not reuse GPT-5.6 Ultra multi-agent scores as single-model apples-to-apples results.

67% gap7 benchmarks
Output price
$50.00 / $30.00
Context
1m / 1.05m
Benchmarks
7 shared
Providers
8 / 2
CodingRAGAgentsLong contextClaude Fable 5 leads SWE-bench Pro

Claude Fable 5 vs Grok 4.5

Pick Grok 4.5 when lower standard-tier API economics ($2/$6 per 1M tokens up to 200K prompt tokens), Grok Build or Cursor distribution, and xAI-first coding launch claims matter most, while accepting first-party-only chart rows for some benchmarks. Pick Claude Fable 5 when Anthropic's stronger sourced agentic coding evidence and adaptive thinking fit the workload and you have verified live Fable 5 access, accepting higher $10/$50 launch pricing and explicit safety-classifier behavior. This is the flagship xAI-versus-Anthropic route; use Sonnet 5 only for balanced-tier product comparisons.

733% gap7 benchmarks
Output price
$50.00 / $6.00
Context
1m / 500k
Benchmarks
7 shared
Providers
8 / 3
CodingRAGAgentsLong contextClaude Fable 5 leads SWE-bench Pro

Claude Fable 5 vs Claude Sonnet 5

Pick Claude Fable 5 when the workload needs Anthropic's highest generally available capability tier and you can accept refusal/fallback behavior plus the $10/M input and $50/M output launch price once access is restored. Pick Claude Sonnet 5 for cost-sensitive production coding, higher throughput, and mainstream API adoption where Sonnet-tier capability is enough.

400% gap10 benchmarks
Output price
$50.00 / $10.00
Context
1m / 1m
Benchmarks
10 shared
Providers
8 / 5
CodingRAGAgentsLong contextClaude Fable 5 leads SWE-bench Verified

Claude Sonnet 4.6 vs Claude Sonnet 5

Claude Sonnet 5 is ~50% cheaper at $2/1M; pay for Claude Sonnet 4.6 only for coding workflow support.

50% gap4 benchmarks
Output price
$15.00 / $10.00
Context
1m / 1m
Benchmarks
4 shared
Providers
6 / 5
CodingRAGAgentsLong contextClaude Sonnet 5 leads SWE-bench Verified

Claude Fable 5 vs Claude Opus 4.8

Pick Claude Fable 5 when you need Anthropic's most capable widely released Mythos-class route and can accept documented refusal and fallback behavior after verifying live access. Keep Claude Opus 4.8 when you need the prior Opus behavior profile, fallback predictability, or a route that may still be easier to reason about for some regulated workflows. Treat the June 9 launch, June 12 suspension, and July 1 redeploy timeline as part of the product decision, not a footnote.

100% gap10 benchmarks
Output price
$50.00 / $25.00
Context
1m / 1m
Benchmarks
10 shared
Providers
8 / 6
CodingRAGAgentsLong contextClaude Fable 5 leads SWE-bench Verified

Claude Fable 5 vs GPT-5.5

On every published agentic coding benchmark, Claude Fable 5 outperforms GPT-5.5 by a wide margin: 80.3% vs 58.6% on SWE-bench Pro (+21.7 pts), 96% vs 82.6% on SWE-bench Verified (Vals.ai), and 85.0% vs 78.7% on OSWorld-Verified computer use. Fable 5 also leads on knowledge-work quality (GDPval-AA ELO: 1932 vs 1769) and agentic legal tasks (13.3% vs 2.1% Legal Agent Benchmark). GPT-5.5 counters with a 93.6% GPQA Diamond score, while Fable 5's GPQA is not published, and notably costs half as much at $5/$30 per 1M tokens versus $10/$50. For pure coding and agentic workflows, Claude Fable 5 is the stronger performer once access is restored. For teams balancing cost with broad reasoning capability, GPT-5.5 is a compelling alternative, especially at roughly half the output price.

67% gap14 benchmarks
Output price
$50.00 / $30.00
Context
1m / 1.05m
Benchmarks
14 shared
Providers
8 / 4
CodingRAGAgentsLong contextClaude Fable 5 leads SWE-bench Verified

Claude Opus 4.7 vs Claude Opus 4.8

Pick Claude Opus 4.8 for higher current agentic coding and computer-use confidence; token pricing is tied on tracked $5/1M input and $25/1M output routes, so keep Claude Opus 4.7 only for already-validated prompts or coding workflow support constraints.

0% gap5 benchmarks
Output price
$25.00 / $25.00
Context
1m / 1m
Benchmarks
5 shared
Providers
6 / 6
CodingRAGAgentsLong contextClaude Opus 4.8 leads SWE-bench Verified

Claude Opus 4.8 vs Gemini 3.5 Pro

Use Claude Opus 4.8 for production agentic coding today: it has tracked provider routes, pricing, and public rows for SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.1, and GPQA Diamond. Track Gemini 3.5 Pro for long-context and multimodal workloads where the 2M-token window may beat Opus 4.8's 1M window, but do not budget or migrate production traffic until Google confirms GA pricing and public API details.

No price gapBenchmark gap
Output price
$25.00 / Unpriced
Context
1m / 2m
Benchmarks
No shared rows
Providers
6 / 0
CodingRAGAgentsLong context

Gemini 3.5 Pro vs GPT-5.5

Pick GPT-5.5 for production decisions today because it has public pricing, multiple provider routes, and benchmark rows in the local data. Keep Gemini 3.5 Pro on the shortlist when the workload is bottlenecked by context length or Google ecosystem routing, but wait for GA pricing, public provider routes, and independent benchmark evidence before replacing GPT-5.5.

No price gapBenchmark gap
Output price
Unpriced / $30.00
Context
2m / 1.05m
Benchmarks
No shared rows
Providers
0 / 4
Long contextVisionCodingRAG

Claude Opus 4.8 vs GPT-5.3-Codex

Pick Claude Opus 4.8 for autonomous repo work, complex multi-file engineering, computer-use agents, and long-context sessions: it leads GPT-5.3-Codex by 12.4 points on SWE-bench Pro and 18.7 points on OSWorld, with 1M context versus 400K. Pick GPT-5.3-Codex for cost-sensitive coding pipelines, OpenAI-native Codex workflows, and terminal automation where its $1.75/M input price and 77.3% Terminal-Bench 2.0 score matter more than the harder agent benchmarks.

79% gap2 benchmarks
Output price
$25.00 / $14.00
Context
1m / 400k
Benchmarks
2 shared
Providers
6 / 3
CodingRAGAgentsLong contextClaude Opus 4.8 leads SWE-bench Verified

Claude Opus 4.8 vs GPT-5.5

Pick Claude Opus 4.8 for coding; GPT-5.5 is better when coding workflow support matters more.

20% gap14 benchmarks
Output price
$25.00 / $30.00
Context
1m / 1.05m
Benchmarks
14 shared
Providers
6 / 4
CodingRAGAgentsLong contextClaude Opus 4.8 leads SWE-bench Verified

DeepSeek V4 Pro vs Unisound U2

Pick DeepSeek V4 Pro for production evaluation today: it has sourced context, pricing routes, and stronger public benchmark coverage. Evaluate Unisound U2 when Chinese sovereign routing, Token Hub access, or long-horizon task-execution positioning matters, but treat its GPQA 87.9 and SWE-bench Verified 75.0 rows as low-confidence vendor claims until independent leaderboards confirm them.

No price gapBenchmark gap
Output price
$0.870 / Unpriced
Context
1m /
Benchmarks
No shared rows
Providers
5 / 1
CodingRAGAgentsLong context

Gemini 3.5 Flash vs GPT-5.5

Gemini 3.5 Flash is safer overall; choose GPT-5.5 when coding workflow support matters.

233% gap12 benchmarks
Output price
$9.00 / $30.00
Context
1.05m / 1.05m
Benchmarks
12 shared
Providers
4 / 4
CodingRAGAgentsLong contextGPT-5.5 leads SWE-bench Verified

DeepSeek V4 Pro vs GLM-5.1

Pick DeepSeek V4 Pro when cost and context length are the bottleneck: it is about 3x cheaper on input at $0.44/M versus $1.40/M, supports a 1M-token window versus 200K, and leads GPQA Diamond 90.1% versus 86.2%. Pick GLM-5.1 when SWE-bench Pro scores and in-model code execution sandbox are the priority, or when Chatbot Arena human preference score is meaningful (1475 versus 1456 on the text arena).

302% gap6 benchmarks
Output price
$0.870 / $3.50
Context
1m / 200k
Benchmarks
6 shared
Providers
5 / 5
CodingRAGAgentsLong contextGLM-5.1 leads SWE-bench Pro

DeepSeek V4 Pro vs Kimi K2.6

Pick DeepSeek V4 Pro for pure code generation, large-codebase analysis, and the lowest per-token cost before its 75% discount expires on 2026-05-31. Pick Kimi K2.6 when your pipeline processes images, screenshots, PDFs, or spreadsheets, or when you need long agent runs with many sequential tool calls.

301% gap12 benchmarks
Output price
$0.870 / $3.49
Context
1m / 262k
Benchmarks
12 shared
Providers
5 / 9
CodingRAGAgentsLong contextDeepSeek V4 Pro leads MMLU PRO

Claude Sonnet 4.6 vs DeepSeek V4 Flash

DeepSeek V4 Flash is ~3233% cheaper at $0.09/1M; pay for Claude Sonnet 4.6 only for coding workflow support.

8233% gap11 benchmarks
Output price
$15.00 / $0.180
Context
1m / 1m
Benchmarks
11 shared
Providers
6 / 5
CodingRAGAgentsLong contextClaude Sonnet 4.6 leads MMLU PRO

Llama 3 70B Instruct vs Llama 3.1 70B Instruct

Pick Llama 3.1 70B Instruct for coding; token pricing is tied, so keep Llama 3 70B Instruct only for already-validated prompts or route constraints.

0% gap2 benchmarks
Output price
$0.400 / $0.400
Context
8k / 128k
Benchmarks
2 shared
Providers
18 / 13
CodingClassificationJSON / Tool useRAGLlama 3.1 70B Instruct leads HumanEval

Popular comparisons

Top model matchups by recent search demand

The matchups buyers actually run before committing to a provider for coding, agents, or build automation.

Top 100
DeepSeek V4 Pro vs GLM-5.1#1 - 17.8K impressionsDeepSeek V4 Pro vs Kimi K2.6#2 - 9.7K impressionsClaude Sonnet 4.6 vs DeepSeek V4 Flash#3 - 5.7K impressionsLlama 3 70B Instruct vs Llama 3.1 70B Instruct#4 - 5.1K impressionsDeepSeek V4 Flash vs Grok 4#5 - 5K impressionsDeepSeek V4 Flash vs Qwen3.6-27B#6 - 4.9K impressionsClaude Sonnet 4.6 vs DeepSeek V4 Pro#7 - 4.5K impressionsGemini 2.5 Flash vs Grok 4#8 - 4.3K impressionsClaude Opus 4.7 vs Kimi K2.6#9 - 4.2K impressionsDeepSeek V4 Flash vs DeepSeek V4 Pro#10 - 4K impressionsDeepSeek V4 Flash vs GLM-5.1#11 - 3.7K impressionsGemini 2.5 Pro vs Grok 4#12 - 3.6K impressionsClaude Sonnet 4.6 vs Kimi K2.6#13 - 3.5K impressionsGrok 3 Mini vs Grok 4#14 - 3.5K impressionsClaude Sonnet 4.6 vs Composer 2.5#15 - 3.4K impressionsQwen3.6-27B vs Qwen3.6-35B-A3B#16 - 3.3K impressionsGPT-5.5 vs o3#17 - 3.3K impressionsDeepSeek V4 Flash vs Kimi K2.6#18 - 3.3K impressionsGLM-5 vs GLM-5.1#19 - 3K impressionsClaude Sonnet 4.6 vs GPT-5.5 Pro#20 - 2.8K impressionsComposer 2.5 vs Grok Build 0.1#21 - 2.7K impressionsDeepSeek V4 Pro vs Grok 4#22 - 2.6K impressionsGemini 2.5 Pro vs o3#23 - 2.6K impressionsDeepSeek V3.1 vs Grok 4#24 - 2.5K impressionsDeepSeek V4 Pro vs Gemini 2.5 Flash#25 - 2.4K impressionsClaude Sonnet 4.6 vs Gemini 3.5 Flash#26 - 2.3K impressionsClaude Opus 4.7 vs DeepSeek V4 Pro#27 - 2.2K impressionsGrok 4 vs Kimi K2.6#28 - 2.1K impressionsDeepSeek V3.1 vs DeepSeek V4 Pro#29 - 2.1K impressionsGrok-3 vs Grok 4#30 - 2.1K impressionsGrok 4 vs Qwen3-Max#31 - 2K impressionsDeepSeek V3 vs Grok 4#32 - 1.9K impressionsClaude Sonnet 4.5 vs DeepSeek V4 Pro#33 - 1.8K impressionsDeepSeek V3.1 vs DeepSeek V4 Flash#34 - 1.8K impressionsGemini 2.5 Flash vs Grok 4.3#35 - 1.8K impressionsClaude Opus 4.7 vs GLM-5.1#36 - 1.8K impressionsDeepSeek R1 vs Kimi K2.6#37 - 1.7K impressionsGPT-5.2 vs GPT-5.5#38 - 1.7K impressionsDeepSeek V4 Flash vs Qwen3.6-35B-A3B#39 - 1.7K impressionsClaude Sonnet 4.6 vs GLM-5.1#40 - 1.6K impressionsGrok 4.3 vs Kimi K2.6#41 - 1.6K impressionsClaude Sonnet 4.5 vs DeepSeek V4 Flash#42 - 1.6K impressionsClaude Opus 4.7 vs Qwen3.6-27B#43 - 1.6K impressionsHunyuan Hy3 Preview vs o3#44 - 1.6K impressionsDeepSeek V4 Flash vs Qwen3-Max#45 - 1.5K impressionsDeepSeek V4 Pro vs Gemini 2.5 Pro#46 - 1.5K impressionsDeepSeek V3 vs Kimi K2.6#47 - 1.5K impressionsDeepSeek V4 Flash vs Gemini 2.5 Flash#48 - 1.5K impressionsQwen3-Max vs Qwen3.6-27B#49 - 1.5K impressionsDeepSeek V4 Pro vs Gemini 3.1 Pro Preview#50 - 1.5K impressionsDeepSeek V4 Flash vs Gemini 2.5 Pro#51 - 1.5K impressionsGLM-5.1 vs GPT-5.5#52 - 1.4K impressionsClaude Opus 4.7 vs Claude Sonnet 4.6#53 - 1.4K impressionsGLM-5.1 vs Xiaomi MiMo-V2.5-Pro#54 - 1.4K impressionsGPT-5.5 vs Kimi K2.6#55 - 1.3K impressionsGemini 3.1 Pro Preview vs Grok 4#56 - 1.3K impressionsDeepSeek R1 vs Grok-3#57 - 1.3K impressionsClaude Mythos Preview vs Grok 4#58 - 1.3K impressionsDeepSeek V4 Pro vs Qwen3.6-27B#59 - 1.3K impressionsClaude Sonnet 4.6 vs Qwen3.6-27B#60 - 1.3K impressionsDeepSeek V4 Pro vs Grok-3#61 - 1.3K impressionsDeepSeek R1 vs DeepSeek V3.1#62 - 1.3K impressionsComposer 2.5 vs Gemini 3.5 Flash#63 - 1.3K impressionsDeepSeek R1 vs Qwen3-235B-A22B#64 - 1.2K impressionsGPT-5.4 vs Kimi K2.6#65 - 1.2K impressionsDeepSeek V4 Pro vs Kimi K2.5#66 - 1.2K impressionsClaude Opus 4.7 vs DeepSeek V4 Flash#67 - 1.2K impressionsGLM-5.1 vs Kimi K2.5#68 - 1.2K impressionsClaude Sonnet 4.6 vs GPT-5.5#69 - 1.2K impressionsGPT-5.5 vs Grok 4#70 - 1.2K impressionsDeepSeek V4 Flash vs Kimi K2.5#71 - 1.1K impressionsClaude Opus 4.6 vs DeepSeek V4 Pro#72 - 1.1K impressionsClaude Haiku 4.5 vs DeepSeek V4 Flash#73 - 1.1K impressionsDeepSeek R1 vs Qwen3-Max#74 - 1.1K impressionsLlama 3.3 70B vs Qwen2.5-72B#75 - 1.1K impressionsClaude Opus 4.7 vs Grok 4#76 - 1.1K impressionsComposer 2.5 vs DeepSeek V4 Pro#77 - 1K impressionsDeepSeek V3.1 vs Kimi K2.6#78 - 1K impressionsDeepSeek V4 Pro vs Qwen3.6-35B-A3B#79 - 1K impressionsDeepSeek V4 Pro vs GPT-5.5#80 - 1K impressionsLlama 3.3 70B Instruct (free) vs Qwen2.5-72B-Instruct#81 - 1K impressionsGemini 3.1 Pro Preview vs Kimi K2.6#82 - 1K impressionsClaude Opus 4.6 vs GLM-5.1#83 - 1K impressionsComposer 2.5 vs GPT-5.5#84 - 1K impressionsQwen3.5-122B-A10B vs Qwen3.6-27B#85 - 995 impressionsComposer 2.5 vs Kimi K2.6#86 - 995 impressionsClaude Sonnet 4.6 vs Kimi K2 Thinking#87 - 972 impressionsClaude Opus 4.7 vs Claude Sonnet 4.5#88 - 907 impressionsDeepSeek R1 Lite vs DeepSeek V4 Pro#89 - 903 impressionsClaude Opus 4.7 vs Composer 2.5#90 - 903 impressionsGPT-5.5 vs o3 Mini#91 - 900 impressionsClaude Opus 4.7 vs Gemini 3.5 Flash#92 - 900 impressionsQwen3.5-35B-A3B vs Qwen3.6-27B#93 - 899 impressionsMiniMax M2 vs MiniMax M3#94 - 880 impressionsDeepSeek V4 Pro vs GLM-5#95 - 872 impressionsLlama 3.3 70B vs Qwen2.5-72B-Instruct#96 - 859 impressionsDeepSeek V4 Pro vs Gemini 3.5 Flash#97 - 842 impressionsGrok 4 vs o3 Mini#98 - 841 impressionsDeepSeek R1 vs Gemini 2.5 Flash#99 - 836 impressionsClaude Opus 4.5 vs GPT-5.5#100 - 825 impressions