LLM Reference

The Best Mainstream LLM APIs, Ranked (2026)

Last refreshed 2026-09-02. Next refresh: weekly.

The best LLM APIs for developers in 2026 — ranked by benchmark quality first, then price. Updated daily with live provider pricing.

Verdict

Use GPT-5.6 Sol for mainstream API work today.

Kimi K3 is the runner-up, 1 point back on GPQA Diamond.

Researched 55d agoWhy this pickMethodology

How we rank

Mainstream API picks now lead with capability: GPQA Diamond first, MMLU as fallback, then price only as the tie-break.

  1. EligibilityChat/completion models that pass the generic API leaderboard filter (no embeddings/rerankers/modality SKUs).
  2. Primary rankingGPQA Diamond, then MMLU when GPQA is missing or tied.
  3. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. Price tie-breakLower tracked input $/1M wins only after the capability scores are exhausted.
  5. PricingRates reflect tracked provider rows — spot/check enterprise tiers separately.
#ModelInput $/1MOutput $/1M
1GPT-5.6 Sol
ReasoningVisionTools

Capability signal: GPQA Diamond 94.6%

$5.00$30.00
2Gemini 3.1 Pro Preview
PreviewVisionTools

Capability signal: GPQA Diamond 94.3%

$2.00$12.00
3Claude Opus 4.7
ReasoningVisionTools

Capability signal: GPQA Diamond 94.2%

$5.00$25.00
4GPT-5.5
ReasoningVisionTools

Capability signal: GPQA Diamond 93.6%

$5.00$30.00
5Claude Opus 4.8
ReasoningVisionTools

Capability signal: GPQA Diamond 93.6%

$5.00$25.00
6GPT-5.5 Pro
ReasoningVisionTools

Capability signal: GPQA Diamond 93.6%

$30.00$180.00
7Kimi K3
ReasoningVisionTools

Capability signal: GPQA Diamond 93.5%

$3.00$15.00
8MiniMax M3
ReasoningVisionTools

Capability signal: GPQA Diamond 92.9%

$0.30$1.20
9Qwen3.8-Max
ReasoningVisionTools

Capability signal: GPQA Diamond 92.6%

$2.00$6.00
10Qwen3.7-Max
ReasoningTools

Capability signal: GPQA Diamond 92.4%

$1.25$3.75
11Gemini 3.5 Flash
ReasoningVisionTools

Capability signal: GPQA Diamond 92.2%

$1.50$9.00
12GPT-5.4
ReasoningVisionTools

Capability signal: GPQA Diamond 92%

$2.50$15.00
13Gemini 3 Pro
VisionTools

Capability signal: GPQA Diamond 91.9%

$1.25$5.00
14Claude Opus 4.6
ReasoningVisionTools

Capability signal: GPQA Diamond 91.3%

$5.00$25.00
15GLM-5.2
ReasoningTools

Capability signal: GPQA Diamond 91.2%

$1.40$4.40
16Kimi K2.6
ReasoningVisionTools

Capability signal: GPQA Diamond 90.5%

$0.73$3.40
17Qwen3.6-Plus
VisionTools

Capability signal: GPQA Diamond 90.4%

$0.33$1.95
18Gemini 3 Flash
PreviewVisionTools

Capability signal: GPQA Diamond 90.4%

$0.50$3.00
19DeepSeek V4 Pro
ReasoningTools

Capability signal: GPQA Diamond 90.1%

$0.43$0.87
20Grok 4.3
ReasoningVisionTools

Capability signal: GPQA Diamond 90.1%

$1.25$2.50

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Qwen3.7-Max

    Alibaba's closed-weight flagship language model, announced at the 2026 Alibaba Cloud Summit (May 20). Scored 56.6 on Artificial Analysis Intelligence Index at launch—highest-ranked Chinese model. 1M-token context with prompt caching (up to 90% discount). Pricing: $2.50/$7.50 per 1M tokens in/out.

    92.4%

    GPQA Diamond

  • #5Gemini 3.5 Flash

    Gemini 3.5 Flash is Google DeepMind's generally available Flash model for sustained frontier-level performance on agentic and coding tasks. It supports multimodal inputs, native thinking, tool and function calling, structured outputs, code execution, search grounding, batch processing, and long contexts up to 1M tokens.

    92.2%

    GPQA Diamond

  • #6GPT-5.4

    GPT-5.4 is OpenAI's flagship frontier reasoning model, released March 5, 2026. It incorporates advances from GPT-5.3-Codex for coding and agentic workflows, and adds 'Thinking' mode with editable reasoning plans. Key capabilities include computer use (navigating interfaces via Playwright), image understanding and generation integration, full-stack web app generation, tool calling, and deep research. Knowledge cutoff is August 31, 2025. Model ID: gpt-5.4.

    92%

    GPQA Diamond