LLM Reference

Best Multimodal LLMs for Vision (2026)

Last refreshed 2026-09-09. Next refresh: weekly.

Best vision and multimodal LLMs in 2026, ranked by image benchmarks. Covers image QA, document understanding, and video analysis.

Verdict

Use Nex-N2.5-mini for vision today.

Muse Spark 1.3 is the runner-up: — vs — on MMMU.

Researched todayWhy this pickMethodology

How we rank

Vision/multimodal leaders rank on standard MMMU, use MathVista only as a comparable near-tie signal, then recency. MMMU Pro is tracked separately as harder multimodal evidence, but models without standard MMMU stay benchmark-pending for this leaderboard.

  1. EligibilityModels flagged `vision` or `multimodal` in seed data, excluding audio, speech, image-generation, and video-generation specialist models.
  2. Primary rankingStandard MMMU score is the primary score. When two scored models are within 1 point and both have MathVista, MathVista breaks the near-tie before release recency.
  3. Benchmark pendingRecent source-backed vision models without tracked standard MMMU stay visible in a separate benchmark-pending section. MMMU Pro rows remain informational unless the methodology is changed to accept that harder benchmark as a proxy.
  4. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. PricingMultimodal pricing often differs by modality — use provider rows for image/video-specific tiers.
#ModelInput $/1MOutput $/1M
1Qwen3.6-Plus
VisionTools

MMMU: 86%

$0.33$1.95
2ByteDance Doubao Seed 2.0 Pro
VisionTools

MMMU: 85.4%

$0.47$2.37
3Qwen3.5-397B-A17B
ReasoningVisionTools

MMMU: 85%

$0.39$2.34
4Gemini 3.5 Flash
ReasoningVisionTools

MMMU: 83.6%

$1.50$9.00
5Claude Sonnet 4.6
ReasoningVisionTools

MMMU: 83.6%

$3.00$15.00
6o3
ReasoningVisionTools

MMMU: 82.9%

$2.00$8.00
7GPT-5.4
ReasoningVisionTools

MMMU: 82.1%

$2.50$15.00
8Qwen3.6 Max Preview
PreviewReasoningVisionTools

MMMU: 82%

$1.04$6.24
9Gemini 2.5 Pro
ReasoningVisionTools

MMMU: 81.7%

$1.25$10.00
10Gemini 3 Pro
VisionTools

MMMU: 81%

$1.25$5.00
11Claude Opus 4.5
ReasoningVisionTools

MMMU: 80.7%

$5.00$25.00
12Gemini 2.5 Flash
VisionTools

MMMU: 79.7%

$0.30$2.50
13Gemini 2.5 Pro Preview 05-06
PreviewVision

MMMU: 79.6%

$1.25$10.00
14Claude Sonnet 4.5
ReasoningVisionTools

MMMU: 77.8%

$3.00$15.00
15Claude Opus 4.6
ReasoningVisionTools

MMMU: 76.5%

$5.00$25.00
16Command A+
ReasoningVisionTools

MMMU: 75.1%

17Claude 3.7 Sonnet
ReasoningVisionTools

MMMU: 75%

$3.00$15.00
18Llama 4 Maverick 17B Instruct FP8
Vision

MMMU: 73.4%

$0.15$0.60
19Llama 4 Scout 17B-16E Instruct
Vision

MMMU: 69.4%

$0.08$0.22
20GPT-4o
VisionTools

MMMU: 69.1%

$2.50$10.00

New models awaiting benchmark coverage

These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.
ModelWhy it is listedStatusTracked price
Nex-N2.5-mini
Tools
Nex-N2.5-mini reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In Free / Out Free
Gemini 3.8 Flash
ToolsCode execution
Gemini 3.8 Flash reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $0.75 / Out $3.75
Muse Spark 1.3
Tools
Muse Spark 1.3 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $1.25 / Out $4.25
Muse Spark 1.3 Contributor
Tools
Muse Spark 1.3 Contributor reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $0.10 / Out $0.20
Qwen3.8-Max-0902
Tools
Qwen3.8-Max-0902 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $2.00 / Out $6.00
Claude Opus 4.8
ToolsCode execution
Claude Opus 4.8 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $25.00
GPT-5.5
ToolsCode execution
GPT-5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $30.00
Claude Opus 4.7
ToolsCode execution
Claude Opus 4.7 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $25.00

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8. Compare it for Coding, Agents, Long context, Vision, and JSON / Tool use.

    MMMU

  • Qwen3.8-Max-0902 is Alibaba's dated API snapshot of Qwen3.8-Max, listed September 2, 2026 on QwenCloud and Model Studio as model ID qwen3.8-max-0902 (first-party alias qwen3.8-max-2026-09-02). Official copy: an upgraded snapshot of qwen3.8-max with stronger coding (engineering-scale / long-horizon autonomous development), collaborative-agent multi-tool orchestration, and refined native vision (chart reasoning, document parsing, multimodal perception). Retains the 1,000,000-token context window, thinking mode, and full tool ecosystem. Input image/text/video, output text; function calling, structured outputs, context cache. International QwenCloud / Model Studio Singapore list price $2 input / $6 output per 1M tokens; implicit cache $0.25; explicit cache create $2.50; explicit cache read $0.17. No first-party 0902 weight drop. Sibling of seeded qwen3.8-max under family qwen3.8. Compare it for Agents and Coding.

    MMMU

  • Anthropic's generally available Mythos-class model for demanding reasoning, long-horizon agents, and ambitious coding, released September 1, 2026. Same underlying model as Claude Mythos 5.1 with cybersecurity and biology safeguards (flagged queries route to Opus 4.8 for cyber and Opus 5 for biology; those reroutes are not billed at Fable prices). Claude API id claude-fable-5-1. List price $10 per 1M input tokens and $50 per 1M output tokens; cache reads $0.25 per 1M (0.025x base input, 75% less than Fable 5's $1.00). 1M-token context, 128K max output, adaptive thinking always on (default effort high), text and image input to text output, vision, tool use. Available to Pro, Max, Team, and Enterprise users and on the Claude Platform natively, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Reliable knowledge cutoff June 2026. Sibling of seeded claude-fable-5 under family claude-fable. Compare it for Agents and Coding.

    MMMU