LLM Reference
activeMathematics

MATH-500

Metric: Score (higher is better)

500-problem subset of the MATH benchmark covering competition mathematics High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Models ranked

7

tracked on this benchmark

Score band

98.0 – 75.2

best → lowest tracked

Snapshot trend

-10.41

Jun 7 → Jun 25 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Kimi K2.5
Version: Thinking mode (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
98.0
2
OLMo 3 32B Think
Version: MATH benchmark (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
96.1
3
Phi-4 Mini Reasoning
Version: From official Microsoft technical report (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
94.6
4
OLMo 3.1 32B Instruct
Version: MATH benchmark (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
93.4
5
LFM2.5 8B A1B
Version: MATH-500 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
88.8
6
Nemotron-Labs TwoTower 30B-A3B Base
Version: 4-shot; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardHarness: Not recordedEvaluator: Not recordedObserved: Jun 25, 2026Confidence: Not recordedSource
80.6
7
Phi-4 Reasoning Vision 15B
Version: MATH-500 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
75.2

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

FAQ

What does the MATH-500 benchmark measure?

500-problem subset of the MATH benchmark covering competition mathematics On this page it lists 7 tracked model variants where higher is better.

Is a higher MATH-500 score always better?

For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.

How current is this MATH-500 data?

This benchmark was last reviewed on May 20, 2026. The tracked score average moved -10.41 points across the last 2 snapshots.

Related benchmarks

Last reviewed: May 20, 2026

Resources