ArXivMath June 2026
MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.
Configurations
2
tracked on this benchmark
Score band
91.3 – 90.8
best → lowest tracked
Snapshot trend
—
need ≥2 snapshots
Leaderboard
Tracked models ranked by % Correct (higher is better).
Configuration: Maximum effort, with tools
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 90.82 no-tools configuration.
Configuration: Maximum effort, no tools
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 91.33 with-tools configuration.
How to read this benchmark
This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.
Trust this score when
- The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
- Cost, token, and step averages are evaluated together with the score for the same configuration.
- You validate the configuration in your own Cursor workload before making a model choice.
Be cautious when
- Rows are collapsed into one score per base model or compared across benchmark versions.
- Small score differences are treated as decisive despite Cursor's variance caveat.
- A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.
FAQ
What does the ArXivMath June 2026 benchmark measure?
MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations. On this page it lists 2 tracked configurations where higher is better.
Is a higher ArXivMath June 2026 score always better?
For this benchmark, higher is better, but Cursor cautions that small differences may not be statistically meaningful. Treat each row as one exact configuration, not a neutral per-model winner, and validate qualified results such as Grok 4.5 separately.
How current is this ArXivMath June 2026 data?
This benchmark was last reviewed on Jul 26, 2026. Re-check the linked model pages for the freshest provider and pricing detail.
Related benchmarks
Last reviewed: Jul 26, 2026