activeMathematics

MATH-500

Metric: Score (higher is better)

500-problem subset of the MATH benchmark covering competition mathematics

Models ranked

7

tracked on this benchmark

Score band

98.0 – 75.2

best → lowest tracked

Snapshot trend

-10.41

Jun 7 → Jun 25 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Kimi K2.5
Version: Thinking mode (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
98.0
2
OLMo 3 32B Think
Version: MATH benchmark (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
96.1
3
Phi-4 Mini Reasoning
Version: From official Microsoft technical report (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
94.6
4
OLMo 3.1 32B Instruct
Version: MATH benchmark (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
93.4
5
LFM2.5 8B A1B
Version: MATH-500 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
88.8
6
Nemotron-Labs TwoTower 30B-A3B Base
Version: 4-shot; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardHarness: Not recordedEvaluator: Not recordedObserved: Jun 25, 2026Confidence: Not recordedSource
80.6
7
Phi-4 Reasoning Vision 15B
Version: MATH-500 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
75.2

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: May 20, 2026

Resources