ArXivMath June 2026
MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations.
Configurations
2
tracked on this benchmark
Score band
91.3 – 90.8
best → lowest tracked
Snapshot trend
—
need ≥2 snapshots
Leaderboard
Tracked models ranked by % Correct (higher is better).
Configuration: Maximum effort, with tools
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 90.82 no-tools configuration.
Configuration: Maximum effort, no tools
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 91.33 with-tools configuration.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
- Cost, token, and step averages are evaluated together with the score for the same configuration.
- You validate the configuration in your own Cursor workload before making a model choice.
Be cautious when
- Rows are collapsed into one score per base model or compared across benchmark versions.
- Small score differences are treated as decisive despite Cursor's variance caveat.
- A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.
Related benchmarks
Last reviewed: Jul 26, 2026