LLM Reference
activeReasoning

MMMU Pro

Metric: Accuracy (higher is better)Introduced: 2024

A harder, more robust extension of MMMU with approximately 3,400 filtered questions and an additional standard setting. Tests graduate-level multimodal reasoning across 30 disciplines while reducing answer shortcuts versus standard MMMU. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Models ranked

51

tracked on this benchmark

Score band

88.3 – 38.5

best → lowest tracked

Snapshot trend

+14.09

Jun 7 → Jul 16 · 1 models

Leaderboard

Tracked models ranked by Accuracy (higher is better).

Compare candidates
#Model variant and provenanceScore
1
GPT-5.5
Version: Vals.ai standardized CoT harnessHarness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource

Notes: DAT-5638: Vals.ai independent harness with chain-of-thought prompting. Pass@1 across about 1,700 MMMU Pro questions, temperature 0, max 8,192 output tokens. Track separately from standard MMMU and do not use the 83.2% tool-augmented/Python variant as the canonical MMMU Pro score.

88.3
2
Gemini 3.5 Flash
Version: Vals.ai standardized CoT harness, 4-option, Pass@1, temp=0Harness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource

Notes: Vals.ai independent evaluation (updated 2026-06-02). Tied #1 with GPT-5.5 at 88.27%; Vals ranks Gemini 3.5 Flash 1st. Google official no-tools score is 83.6% (https://deepmind.google/models/gemini/) - lower due to different setup. Use Vals.ai as canonical (consistent with existing gpt-5.5 row policy).

88.3
3
Gemini 3.1 Pro Preview
Version: Vals.ai standardized CoT harness, 4-option, Pass@1, temp=0Harness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource

Notes: Vals.ai independent evaluation. Vals labels this 'Gemini 3.1 Pro Preview (02/26)' (February 2026 snapshot). Google DeepMind official model card shows 80.5% (Thinking High, no tools) - the Vals CoT premium accounts for the gap.

88.2
4
Gemini 3 Flash
Version: Vals.ai standardized CoT harness, 4-option, Pass@1, temp=0Harness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource

Notes: Vals.ai independent evaluation. Vals labels 'Gemini 3 Flash (12/25)' (December 2025 snapshot). LLM-Stats shows 81.2% (likely official no-tools). Use Vals.ai for consistency with other Vals rows.

87.6
5
GPT-5.6 Sol
Version: MMMU Pro; no tools; max reasoningHarness: Not recordedEvaluator: Not recordedObserved: Jul 9, 2026Confidence: Not recordedSource

Notes: Official GPT-5.6 GA launch no-tools MMMU Pro row. GA launch also cites 84.6% with tools; kept separate in notes because modelBenchmark is keyed by benchmarkSlug only.

83.0
6
Kimi K3
Version: MMMU-Pro; official protocol; three-run mean; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: MMMU-Pro, official protocol, three-run mean. Harness/evaluator: official MMMU-Pro protocol; evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 81.6.

81.6
7
GPT-5.4
Version: official OpenAI, without tool useHarness: Not recordedEvaluator: Not recordedObserved: Apr 1, 2026Confidence: Not recordedSource

Notes: OpenAI official, explicitly 'without tool use'. LLM-Stats confirms 81.2%. With-tools figure not published separately for this model. Vals.ai has not evaluated GPT-5.4 as of 2026-06-07.

81.2
8
Gemini 3 Pro
Version: official Google DeepMind model card, Thinking High, no toolsHarness: Not recordedEvaluator: Not recordedObserved: Nov 18, 2025Confidence: Not recordedSource

Notes: Official score from Google DeepMind Gemini 3.1 Pro model card comparison table (Gemini 3 Pro Thinking High = 81.0%). Confirmed by Google I/O launch announcement and LLM-Stats (81.0%). Vals.ai shows 87.51% via CoT harness - higher but consistent with Vals premium on other models. NOTE: Unlike other Gemini models, no Vals row available at the required comparable level. Using official score here. Consider backfilling with Vals score if/when Vals adds a comparable Vals row for Gemini 3 Pro.

81.0
9
Muse Spark
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Meta Muse Spark. LLM-Stats only source; no official Meta MMMU-Pro announcement found. Confidence: medium.

80.4
10
Kimi K2.6
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Moonshot AI Kimi K2.6. LLM-Stats only source. Confidence: medium.

80.1
11
GPT-5.2
Version: official OpenAI, without tool useHarness: Not recordedEvaluator: Not recordedObserved: Jan 1, 2026Confidence: Not recordedSource

Notes: OpenAI official no-tools score (79.5%). With-tools=80.4%, with Python=86.5%. Using no-tools for comparability with other official scores. LLM-Stats confirms 79.5%.

79.5
12
Qwen3.6-Plus
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. LLM-Stats only source.

78.8
13
Kimi K2.5
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. LLM-Stats only source.

78.5
14
GPT-5
Version: thinking mode (highest across thinking levels)Harness: Not recordedEvaluator: Not recordedObserved: May 14, 2025Confidence: Not recordedSource

Notes: GPT-5 without thinking/tools scores 62.7%; thinking mode scores 78.4%. Per CLAUDE.md policy, use highest score across thinking levels. LLM-Stats confirms 78.4%.

78.4
15
MiniMax M3
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. LLM-Stats only source.

78.1
16
Claude Opus 4.6
Version: official Anthropic system card, adaptive thinking, max effort, with image cropping tool, avg 5 runsHarness: Not recordedEvaluator: Not recordedObserved: Feb 1, 2026Confidence: Not recordedSource

Notes: Anthropic official score with tools (77.3%). No-tools score is 73.9% from same card. Using with-tools as it represents full model capability. LLM-Stats confirms 77.3%.

77.3
17
Gemma 4 31B
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium.

76.9
18
Qwen3.5-122B-A10B
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium.

76.9
19
Gemini 3.1 Flash-Lite
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium.

76.8
20
GPT-5.4 Mini
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. BenchLM.ai shows 76.6% confirming LLM-Stats.

76.6
21
o3
Version: LLM-Stats aggregator (official no-tools or equivalent)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: No official OpenAI MMMU-Pro score found for o3. OpenAI o3 launch announced MMMU at 82.9% (standard MMMU, not MMMU-Pro). LLM-Stats is the best available source for MMMU-Pro. Confidence: medium.

76.4
22
GPT-5.5 Instant
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium.

76.0
23
Qwen3.6-27B
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. BenchLM.ai confirms 75.8%.

75.8
24
Claude Sonnet 4.6
Version: official Anthropic system card, adaptive thinking, max effort, with image cropping toolHarness: Not recordedEvaluator: Not recordedObserved: Feb 17, 2026Confidence: Not recordedSource

Notes: Anthropic official with-tools score (75.6%). No-tools is 74.5%. LLM-Stats confirms 75.6%.

75.6
25
Qwen3.6-35B-A3B
Version: LLM-Stats aggregatorHarness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource

Notes: Confidence: medium. BenchLM.ai confirms 75.3%.

75.3
Showing the top 25 of 51 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

FAQ

What does the MMMU Pro benchmark measure?

A harder, more robust extension of MMMU with approximately 3,400 filtered questions and an additional standard setting. Tests graduate-level multimodal reasoning across 30 disciplines while reducing answer shortcuts versus standard MMMU. On this page it lists 51 tracked model variants where higher is better.

Is a higher MMMU Pro score always better?

For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.

How current is this MMMU Pro data?

This benchmark was last reviewed on Jun 5, 2026. The tracked score average moved +14.09 points across the last 3 snapshots.

Related benchmarks

Last reviewed: Jun 5, 2026