activeMathematics

AIME 2024

Metric: Score (higher is better)

American Invitational Mathematics Examination 2024 problems

Models ranked

9

tracked on this benchmark

Score band

96.7 – 57.5

best → lowest tracked

Snapshot trend

-12.92

Apr 29 → Jun 7 · 4 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
o3
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 16, 2025Confidence: Not recordedSource
96.7
2
Gemini 2.5 Pro
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Mar 25, 2025Confidence: Not recordedSource
92.0
3
DeepSeek R1 0528
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2025Confidence: Not recordedSource
91.4
4
Qwen3-Coder-Next
Version: From official technical report arXiv:2603 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
89.0
5
o3 Mini
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Jan 31, 2025Confidence: Not recordedSource
87.3
6
Qwen3-235B-A22B
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 29, 2025Confidence: Not recordedSource
85.7
7
OLMo 3 32B Think
Version: AIME 2024 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
76.8
8
OLMo 3.1 32B Instruct
Version: AIME 2024 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
67.8
9
Phi-4 Mini Reasoning
Version: From official Microsoft technical report, Table 3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
57.5

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: May 20, 2026

Resources