LLM Reference
activeReasoning

ArXivMath June 2026

Metric: % Correct (higher is better)Introduced: 2026

MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Configurations

2

tracked on this benchmark

Score band

91.3 – 90.8

best → lowest tracked

Snapshot trend

need ≥2 snapshots

Leaderboard

Tracked models ranked by % Correct (higher is better).

Compare candidates
#Published configuration and provenanceScore
1
Claude Opus 5

Configuration: Maximum effort, with tools

Version: June 2026Harness: 49-problem June 2026 set; four runs per problemEvaluator: AnthropicObserved: Jul 24, 2026Confidence: confirmedSource

Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 90.82 no-tools configuration.

91.3
2
Claude Opus 5

Configuration: Maximum effort, no tools

Version: June 2026Harness: 49-problem June 2026 set; four runs per problemEvaluator: AnthropicObserved: Jul 24, 2026Confidence: confirmedSource

Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 91.33 with-tools configuration.

90.8

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
  • Cost, token, and step averages are evaluated together with the score for the same configuration.
  • You validate the configuration in your own Cursor workload before making a model choice.

Be cautious when

  • Rows are collapsed into one score per base model or compared across benchmark versions.
  • Small score differences are treated as decisive despite Cursor's variance caveat.
  • A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.

FAQ

What does the ArXivMath June 2026 benchmark measure?

MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations. On this page it lists 2 tracked configurations where higher is better.

Is a higher ArXivMath June 2026 score always better?

For this benchmark, higher is better, but Cursor cautions that small differences may not be statistically meaningful. Treat each row as one exact configuration, not a neutral per-model winner, and validate qualified results such as Grok 4.5 separately.

How current is this ArXivMath June 2026 data?

This benchmark was last reviewed on Jul 26, 2026. Re-check the linked model pages for the freshest provider and pricing detail.

Related benchmarks

Last reviewed: Jul 26, 2026

Resources