MMMU Pro
A harder, more robust extension of MMMU with approximately 3,400 filtered questions and an additional standard setting. Tests graduate-level multimodal reasoning across 30 disciplines while reducing answer shortcuts versus standard MMMU.
Models ranked
51
tracked on this benchmark
Score band
88.3 – 38.5
best → lowest tracked
Snapshot trend
+14.09
Jun 7 → Jul 16 · 1 models
Leaderboard
Tracked models ranked by Accuracy (higher is better).
Notes: DAT-5638: Vals.ai independent harness with chain-of-thought prompting. Pass@1 across about 1,700 MMMU Pro questions, temperature 0, max 8,192 output tokens. Track separately from standard MMMU and do not use the 83.2% tool-augmented/Python variant as the canonical MMMU Pro score.
Notes: Vals.ai independent evaluation (updated 2026-06-02). Tied #1 with GPT-5.5 at 88.27%; Vals ranks Gemini 3.5 Flash 1st. Google official no-tools score is 83.6% (https://deepmind.google/models/gemini/) - lower due to different setup. Use Vals.ai as canonical (consistent with existing gpt-5.5 row policy).
Notes: Vals.ai independent evaluation. Vals labels this 'Gemini 3.1 Pro Preview (02/26)' (February 2026 snapshot). Google DeepMind official model card shows 80.5% (Thinking High, no tools) - the Vals CoT premium accounts for the gap.
Notes: Vals.ai independent evaluation. Vals labels 'Gemini 3 Flash (12/25)' (December 2025 snapshot). LLM-Stats shows 81.2% (likely official no-tools). Use Vals.ai for consistency with other Vals rows.
Notes: Official GPT-5.6 GA launch no-tools MMMU Pro row. GA launch also cites 84.6% with tools; kept separate in notes because modelBenchmark is keyed by benchmarkSlug only.
Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: MMMU-Pro, official protocol, three-run mean. Harness/evaluator: official MMMU-Pro protocol; evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 81.6.
Notes: OpenAI official, explicitly 'without tool use'. LLM-Stats confirms 81.2%. With-tools figure not published separately for this model. Vals.ai has not evaluated GPT-5.4 as of 2026-06-07.
Notes: Official score from Google DeepMind Gemini 3.1 Pro model card comparison table (Gemini 3 Pro Thinking High = 81.0%). Confirmed by Google I/O launch announcement and LLM-Stats (81.0%). Vals.ai shows 87.51% via CoT harness - higher but consistent with Vals premium on other models. NOTE: Unlike other Gemini models, no Vals row available at the required comparable level. Using official score here. Consider backfilling with Vals score if/when Vals adds a comparable Vals row for Gemini 3 Pro.
Notes: Meta Muse Spark. LLM-Stats only source; no official Meta MMMU-Pro announcement found. Confidence: medium.
Notes: Moonshot AI Kimi K2.6. LLM-Stats only source. Confidence: medium.
Notes: OpenAI official no-tools score (79.5%). With-tools=80.4%, with Python=86.5%. Using no-tools for comparability with other official scores. LLM-Stats confirms 79.5%.
Notes: Confidence: medium. LLM-Stats only source.
Notes: Confidence: medium. LLM-Stats only source.
Notes: GPT-5 without thinking/tools scores 62.7%; thinking mode scores 78.4%. Per CLAUDE.md policy, use highest score across thinking levels. LLM-Stats confirms 78.4%.
Notes: Confidence: medium. LLM-Stats only source.
Notes: Anthropic official score with tools (77.3%). No-tools score is 73.9% from same card. Using with-tools as it represents full model capability. LLM-Stats confirms 77.3%.
Notes: Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium. BenchLM.ai shows 76.6% confirming LLM-Stats.
Notes: No official OpenAI MMMU-Pro score found for o3. OpenAI o3 launch announced MMMU at 82.9% (standard MMMU, not MMMU-Pro). LLM-Stats is the best available source for MMMU-Pro. Confidence: medium.
Notes: Confidence: medium.
Notes: Confidence: medium. BenchLM.ai confirms 75.8%.
Notes: Anthropic official with-tools score (75.6%). No-tools is 74.5%. LLM-Stats confirms 75.6%.
Notes: Confidence: medium. BenchLM.ai confirms 75.3%.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Related benchmarks
Last reviewed: Jun 5, 2026