Gemini 3 Pro
- MMMU
- 81%
- Output (from)
- $5.00 / 1M
Last refreshed 2026-09-30. Next refresh: weekly.
Best vision and multimodal LLMs in 2026, ranked by image benchmarks. Covers image QA, document understanding, and video analysis.
Verdict
GPT-6.1 Sol is the runner-up: 81% vs — on MMMU.
Vision/multimodal leaders rank on standard MMMU, use MathVista only as a comparable near-tie signal, then recency. MMMU Pro is tracked separately as harder multimodal evidence, but models without standard MMMU stay benchmark-pending for this leaderboard.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Qwen3.6-Plus VisionTools MMMU: 86% | $0.33 | $1.95 | |
| 2 | ByteDance Doubao Seed 2.0 Pro VisionTools MMMU: 85.4% | $0.47 | $2.37 | |
| 3 | Qwen3.5-397B-A17B ReasoningVisionTools MMMU: 85% | $0.39 | $2.34 | |
| 4 | Gemini 3.5 Flash ReasoningVisionTools MMMU: 83.6% | $1.50 | $9.00 | |
| 5 | Claude Sonnet 4.6 ReasoningVisionTools MMMU: 83.6% | $3.00 | $15.00 | |
| 6 | o3 ReasoningVisionTools MMMU: 82.9% | $2.00 | $8.00 | |
| 7 | GPT-5.4 ReasoningVisionTools MMMU: 82.1% | $2.50 | $15.00 | |
| 8 | Qwen3.6 Max Preview PreviewReasoningVisionTools MMMU: 82% | $1.04 | $6.24 | |
| 9 | Gemini 2.5 Pro ReasoningVisionTools MMMU: 81.7% | $1.25 | $10.00 | |
| 10 | Gemini 3 Pro VisionTools MMMU: 81% | $1.25 | $5.00 | |
| 11 | Claude Opus 4.5 ReasoningVisionTools MMMU: 80.7% | $5.00 | $25.00 | |
| 12 | Gemini 2.5 Flash VisionTools MMMU: 79.7% | $0.30 | $2.50 | |
| 13 | Gemini 2.5 Pro Preview 05-06 PreviewVision MMMU: 79.6% | $1.25 | $10.00 | |
| 14 | Claude Sonnet 4.5 ReasoningVisionTools MMMU: 77.8% | $3.00 | $15.00 | |
| 15 | Claude Opus 4.6 ReasoningVisionTools MMMU: 76.5% | $5.00 | $25.00 | |
| 16 | Command A+ ReasoningVisionTools MMMU: 75.1% | — | — | |
| 17 | Claude 3.7 Sonnet ReasoningVisionTools MMMU: 75% | $3.00 | $15.00 | |
| 18 | Llama 4 Maverick 17B Instruct FP8 Vision MMMU: 73.4% | $0.15 | $0.60 | |
| 19 | Llama 4 Scout 17B-16E Instruct Vision MMMU: 69.4% | $0.08 | $0.22 | |
| 20 | GPT-4o VisionTools MMMU: 69.1% | $2.50 | $10.00 |
| Model | Why it is listed | Status | Tracked price |
|---|---|---|---|
| GPT-6.1 Sol ToolsCode execution | GPT-6.1 Sol reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $2.00 / Out $10.00 |
| Claude Sonnet 5.5 Tools | Claude Sonnet 5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $2.00 / Out $10.00 |
| Perceptron Mk1.5 Tools | Perceptron Mk1.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $0.15 / Out $1.50 |
| Ember-1 Tools | Ember-1 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $3.00 / Out $15.00 |
| Claude Opus 5.5 Tools | Claude Opus 5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $4.00 / Out $20.00 |
| Claude Opus 4.8 ToolsCode execution | Claude Opus 4.8 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $5.00 / Out $25.00 |
| GPT-5.5 ToolsCode execution | GPT-5.5 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $5.00 / Out $30.00 |
| Claude Opus 4.7 ToolsCode execution | Claude Opus 4.7 reports source-backed vision or multimodal capability; keep it separate from the scored vision ranking until standard MMMU data lands. | Benchmark pending No tracked standard MMMU score yet. | In $5.00 / Out $25.00 |
MiniMax M3.1 Flash Preview (API id MiniMax-M3.1-Flash-Preview) is MiniMax's faster/lighter M3-series multimodal coding model announced for MiniMax Code on 2026-09-27 and made live on Token Plan the same day. First-party docs: 1,000,000-token context; text+image+video inputs to text; agentic reasoning, tool use, and coding; thinking always on with tunable effort low|medium|high|xhigh|max (default max; cannot disable — 400 'requires adaptive thinking'). Availability for now is Token Plan + MiniMax Code only — NOT on pay-as-you-go (paygo price table lists M3/M2.x only; no M3.1 Flash $/MTok row). No public Hugging Face weights/model card for the Flash Preview SKU. Closed-source hosted API — licenseSlug proprietary / weights notReleased.
—
MMMU
Perceptron Mk1.5 is Perceptron's embodied vision-language model released September 25, 2026 (first-party blog). Successor to Mk1 in the Perceptron Mk family: native audio in, video object tracking, web search, sub-agent/tool calls, and stronger visual reasoning; blog cites ~2-5x faster end-to-end vs Mk1. Accepts text, images, video, and audio; emits text plus spatial/temporal annotations (points, boxes, polygons, clips, object tracks). First-party docs: Model ID perceptron-mk1.5; context window 36,864 tokens; max output 8,192; pricing $0.15/M input, $1.50/M output, $0.0375/M cached input. Available via Perceptron Platform + SDK and OpenRouter (perceptron/perceptron-mk1.5). Closed-source hosted API under a proprietary license.
—
MMMU
Ember-1 is Fireworks Research's specialized foundation model announced September 23, 2026, the first model in the Ember series. Built on Kimi K3 and trained by Fireworks (50+ training experiments; Fireworks Serverless Training) to shorten reasoning traces (~40% fewer tokens while holding quality on Fireworks' evaluations and live A/B tests). Fireworks markets Ember-1 as its own model, a Research Preview on Serverless alongside base Kimi K3, and catalogs it as Fireworks Ember rather than a Kimi family member. Fireworks model page: API path accounts/fireworks/models/ember-1; MoE; parameters 2.78T; context length 1040k tokens; image input supported; function calling supported; fine-tuning not supported; serverless supported; kind base model; created 2026-09-22; state Ready; provider Fireworks. Serverless list is $3.00 input / $0.30 cached input / $15.00 output per 1M tokens. No first-party public weights or open license is stated.
—
MMMU