LLM Reference
activeReasoning

ArXivMath June 2026

Metric: % Correct (higher is better)Introduced: 2026

MathArena's June 2026 ArXivMath set is a dynamic 49-problem evaluation derived from recent arXiv mathematics papers. It evaluates final-answer accuracy on newly published problems; tool-assisted and no-tools runs remain separate model-score configurations.

Configurations

2

tracked on this benchmark

Score band

91.3 – 90.8

best → lowest tracked

Snapshot trend

need ≥2 snapshots

Leaderboard

Tracked models ranked by % Correct (higher is better).

Compare candidates
#Published configuration and provenanceScore
1
Claude Opus 5

Configuration: Maximum effort, with tools

Version: June 2026Harness: 49-problem June 2026 set; four runs per problemEvaluator: AnthropicObserved: Jul 24, 2026Confidence: confirmedSource

Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 90.82 no-tools configuration.

91.3
2
Claude Opus 5

Configuration: Maximum effort, no tools

Version: June 2026Harness: 49-problem June 2026 set; four runs per problemEvaluator: AnthropicObserved: Jul 24, 2026Confidence: confirmedSource

Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 91.33 with-tools configuration.

90.8

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
  • Cost, token, and step averages are evaluated together with the score for the same configuration.
  • You validate the configuration in your own Cursor workload before making a model choice.

Be cautious when

  • Rows are collapsed into one score per base model or compared across benchmark versions.
  • Small score differences are treated as decisive despite Cursor's variance caveat.
  • A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.

Related benchmarks

Last reviewed: Jul 26, 2026

Resources