activeMathematics
AIME 2024
Metric: Score (higher is better)
American Invitational Mathematics Examination 2024 problems
Models ranked
9
tracked on this benchmark
Score band
96.7 – 57.5
best → lowest tracked
Snapshot trend
-12.92
Apr 29 → Jun 7 · 4 models
Leaderboard
Tracked models ranked by Score (higher is better).
#Model variant and provenanceRelative to leaderScore
1
o3
96.7Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 16, 2025Confidence: Not recordedSource
2
Gemini 2.5 Pro
92.0Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Mar 25, 2025Confidence: Not recordedSource
3
DeepSeek R1 0528
91.4Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2025Confidence: Not recordedSource
4
Qwen3-Coder-Next
89.0Version: From official technical report arXiv:2603 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
5
o3 Mini
87.3Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Jan 31, 2025Confidence: Not recordedSource
6
Qwen3-235B-A22B
85.7Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 29, 2025Confidence: Not recordedSource
7
OLMo 3 32B Think
76.8Version: AIME 2024 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
8
OLMo 3.1 32B Instruct
67.8Version: AIME 2024 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
9
Phi-4 Mini Reasoning
57.5Version: From official Microsoft technical report, Table 3 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Related benchmarks
Last reviewed: May 20, 2026