MMMU Pro
A harder, more robust extension of MMMU with approximately 3,400 filtered questions and an additional standard setting. Tests graduate-level multimodal reasoning across 30 disciplines while reducing answer shortcuts versus standard MMMU. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.
Models ranked
51
tracked on this benchmark
Score band
88.3 – 38.5
best → lowest tracked
Snapshot trend
+14.09
Jun 7 → Jul 16 · 1 models
Leaderboard
Tracked models ranked by Accuracy (higher is better).
Notes: DAT-5638: Vals.ai independent harness with chain-of-thought prompting. Pass@1 across about 1,700 MMMU Pro questions, temperature 0, max 8,192 output tokens. Track separately from standard MMMU and do not use the 83.2% tool-augmented/Python variant as the canonical MMMU Pro score.
Notes: Vals.ai independent evaluation (updated 2026-06-02). Tied #1 with GPT-5.5 at 88.27%; Vals ranks Gemini 3.5 Flash 1st. Google official no-tools score is 83.6% (https://deepmind.google/models/gemini/) - lower due to different setup. Use Vals.ai as canonical (consistent with existing gpt-5.5 row policy).
Notes: Vals.ai independent evaluation. Vals labels this 'Gemini 3.1 Pro Preview (02/26)' (February 2026 snapshot). Google DeepMind official model card shows 80.5% (Thinking High, no tools) - the Vals CoT premium accounts for the gap.
Notes: Vals.ai independent evaluation. Vals labels 'Gemini 3 Flash (12/25)' (December 2025 snapshot). LLM-Stats shows 81.2% (likely official no-tools). Use Vals.ai for consistency with other Vals rows.
Notes: Official GPT-5.6 GA launch no-tools MMMU Pro row. GA launch also cites 84.6% with tools; kept separate in notes because modelBenchmark is keyed by benchmarkSlug only.
Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: MMMU-Pro, official protocol, three-run mean. Harness/evaluator: official MMMU-Pro protocol; evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 81.6.
Notes: OpenAI official, explicitly 'without tool use'. LLM-Stats confirms 81.2%. With-tools figure not published separately for this model. Vals.ai has not evaluated GPT-5.4 as of 2026-06-07.
Notes: Official score from Google DeepMind Gemini 3.1 Pro model card comparison table (Gemini 3 Pro Thinking High = 81.0%). Confirmed by Google I/O launch announcement and LLM-Stats (81.0%). Vals.ai shows 87.51% via CoT harness - higher but consistent with Vals premium on other models. NOTE: Unlike other Gemini models, no Vals row available at the required comparable level. Using official score here. Consider backfilling with Vals score if/when Vals adds a comparable Vals row for Gemini 3 Pro.
Notes: Meta Muse Spark. LLM-Stats only source; no official Meta MMMU-Pro announcement found. Confidence: medium.
Notes: Moonshot AI Kimi K2.6. LLM-Stats only source. Confidence: medium.
Notes: OpenAI official no-tools score (79.5%). With-tools=80.4%, with Python=86.5%. Using no-tools for comparability with other official scores. LLM-Stats confirms 79.5%.
Notes: Confidence: medium. LLM-Stats only source.
Notes: Confidence: medium. LLM-Stats only source.
Notes: GPT-5 without thinking/tools scores 62.7%; thinking mode scores 78.4%. Per CLAUDE.md policy, use highest score across thinking levels. LLM-Stats confirms 78.4%.
Notes: Confidence: medium. LLM-Stats only source.
Notes: Anthropic official score with tools (77.3%). No-tools score is 73.9% from same card. Using with-tools as it represents full model capability. LLM-Stats confirms 77.3%.
Notes: Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium. BenchLM.ai shows 76.6% confirming LLM-Stats.
Notes: No official OpenAI MMMU-Pro score found for o3. OpenAI o3 launch announced MMMU at 82.9% (standard MMMU, not MMMU-Pro). LLM-Stats is the best available source for MMMU-Pro. Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium. BenchLM.ai confirms 75.8%.
Notes: Anthropic official with-tools score (75.6%). No-tools is 74.5%. LLM-Stats confirms 75.6%.
Notes: Confidence: medium. BenchLM.ai confirms 75.3%.
How to read this benchmark
This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
FAQ
What does the MMMU Pro benchmark measure?
A harder, more robust extension of MMMU with approximately 3,400 filtered questions and an additional standard setting. Tests graduate-level multimodal reasoning across 30 disciplines while reducing answer shortcuts versus standard MMMU. On this page it lists 51 tracked model variants where higher is better.
Is a higher MMMU Pro score always better?
For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.
How current is this MMMU Pro data?
This benchmark was last reviewed on Jun 5, 2026. The tracked score average moved +14.09 points across the last 3 snapshots.
Related benchmarks
Last reviewed: Jun 5, 2026