Best Multimodal LLMs for Vision (2026)

Last refreshed 2026-09-30. Next refresh: weekly.

Best vision and multimodal LLMs in 2026, ranked by image benchmarks. Covers image QA, document understanding, and video analysis.

Verdict

Use Gemini 3 Pro for vision today.

GPT-6.1 Sol is the runner-up: 81% vs — on MMMU.

Researched 9d agoWhy this pickMethodology

How we rank

Vision/multimodal leaders rank on standard MMMU, use MathVista only as a comparable near-tie signal, then recency. MMMU Pro is tracked separately as harder multimodal evidence, but models without standard MMMU stay benchmark-pending for this leaderboard.

  1. Eligibility — Models flagged `vision` or `multimodal` in seed data, excluding audio, speech, image-generation, and video-generation specialist models.
  2. Primary ranking — Standard MMMU score is the primary score. When two scored models are within 1 point and both have MathVista, MathVista breaks the near-tie before release recency.
  3. Benchmark pending — Recent source-backed vision models without tracked standard MMMU stay visible in a separate benchmark-pending section. MMMU Pro rows remain informational unless the methodology is changed to accept that harder benchmark as a proxy.
  4. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. Pricing — Multimodal pricing often differs by modality — use provider rows for image/video-specific tiers.
#ModelInput $/1MOutput $/1M
1Qwen3.6-Plus
VisionTools

MMMU: 86%

$0.33$1.95
2ByteDance Doubao Seed 2.0 Pro
VisionTools

MMMU: 85.4%

$0.47$2.37
3Qwen3.5-397B-A17B
ReasoningVisionTools

MMMU: 85%

$0.39$2.34
4Gemini 3.5 Flash
ReasoningVisionTools

MMMU: 83.6%

$1.50$9.00
5Claude Sonnet 4.6
ReasoningVisionTools

MMMU: 83.6%

$3.00$15.00
6o3
ReasoningVisionTools

MMMU: 82.9%

$2.00$8.00
7GPT-5.4
ReasoningVisionTools

MMMU: 82.1%

$2.50$15.00
8Qwen3.6 Max Preview
PreviewReasoningVisionTools

MMMU: 82%

$1.04$6.24
9Gemini 2.5 Pro
ReasoningVisionTools

MMMU: 81.7%

$1.25$10.00
10Gemini 3 Pro
VisionTools

MMMU: 81%

$1.25$5.00
11Claude Opus 4.5
ReasoningVisionTools

MMMU: 80.7%

$5.00$25.00
12Gemini 2.5 Flash
VisionTools

MMMU: 79.7%

$0.30$2.50
13Gemini 2.5 Pro Preview 05-06
PreviewVision

MMMU: 79.6%

$1.25$10.00
14Claude Sonnet 4.5
ReasoningVisionTools

MMMU: 77.8%

$3.00$15.00
15Claude Opus 4.6
ReasoningVisionTools

MMMU: 76.5%

$5.00$25.00
16Command A+
ReasoningVisionTools

MMMU: 75.1%

——
17Claude 3.7 Sonnet
ReasoningVisionTools

MMMU: 75%

$3.00$15.00
18Llama 4 Maverick 17B Instruct FP8
Vision

MMMU: 73.4%

$0.15$0.60
19Llama 4 Scout 17B-16E Instruct
Vision

MMMU: 69.4%

$0.08$0.22
20GPT-4o
VisionTools

MMMU: 69.1%

$2.50$10.00

New models awaiting benchmark coverage

These source-backed rows qualify for this task page, but they are not scored leaderboard picks until the category benchmark data exists.
ModelWhy it is listedStatusTracked price
GPT-6.1 Sol
ToolsCode execution
GPT-6.1 Sol reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $2.00 / Out $10.00
Claude Sonnet 5.5
Tools
Claude Sonnet 5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $2.00 / Out $10.00
Perceptron Mk1.5
Tools
Perceptron Mk1.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $0.15 / Out $1.50
Ember-1
Tools
Ember-1 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $3.00 / Out $15.00
Claude Opus 5.5
Tools
Claude Opus 5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $4.00 / Out $20.00
Claude Opus 4.8
ToolsCode execution
Claude Opus 4.8 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $25.00
GPT-5.5
ToolsCode execution
GPT-5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $30.00
Claude Opus 4.7
ToolsCode execution
Claude Opus 4.7 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands.Benchmark pending

No tracked standard MMMU score yet.

In $5.00 / Out $25.00

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • MiniMax M3.1 Flash Preview (API id MiniMax-M3.1-Flash-Preview) is MiniMax's faster/lighter M3-series multimodal coding model announced for MiniMax Code on 2026-09-27 and made live on Token Plan the same day. First-party docs: 1,000,000-token context; text+image+video inputs to text; agentic reasoning, tool use, and coding; thinking always on with tunable effort low|medium|high|xhigh|max (default max; cannot disable — 400 'requires adaptive thinking'). Availability for now is Token Plan + MiniMax Code only — NOT on pay-as-you-go (paygo price table lists M3/M2.x only; no M3.1 Flash $/MTok row). No public Hugging Face weights/model card for the Flash Preview SKU. Closed-source hosted API — licenseSlug proprietary / weights notReleased.

    —

    MMMU

  • Perceptron Mk1.5 is Perceptron's embodied vision-language model released September 25, 2026 (first-party blog). Successor to Mk1 in the Perceptron Mk family: native audio in, video object tracking, web search, sub-agent/tool calls, and stronger visual reasoning; blog cites ~2-5x faster end-to-end vs Mk1. Accepts text, images, video, and audio; emits text plus spatial/temporal annotations (points, boxes, polygons, clips, object tracks). First-party docs: Model ID perceptron-mk1.5; context window 36,864 tokens; max output 8,192; pricing $0.15/M input, $1.50/M output, $0.0375/M cached input. Available via Perceptron Platform + SDK and OpenRouter (perceptron/perceptron-mk1.5). Closed-source hosted API under a proprietary license.

    —

    MMMU

  • Ember-1 is Fireworks Research's specialized foundation model announced September 23, 2026, the first model in the Ember series. Built on Kimi K3 and trained by Fireworks (50+ training experiments; Fireworks Serverless Training) to shorten reasoning traces (~40% fewer tokens while holding quality on Fireworks' evaluations and live A/B tests). Fireworks markets Ember-1 as its own model, a Research Preview on Serverless alongside base Kimi K3, and catalogs it as Fireworks Ember rather than a Kimi family member. Fireworks model page: API path accounts/fireworks/models/ember-1; MoE; parameters 2.78T; context length 1040k tokens; image input supported; function calling supported; fine-tuning not supported; serverless supported; kind base model; created 2026-09-22; state Ready; provider Fireworks. Serverless list is $3.00 input / $0.30 cached input / $15.00 output per 1M tokens. No first-party public weights or open license is stated.

    —

    MMMU