The 42 rows below are LLMReference's July 18, 2026 snapshot of CursorBench 3.2. CursorBench is a vendor-run evaluation of IDE-native, multi-file coding-agent workflows, not standalone base models, and its public harness is not independently reproducible. Each reasoning or effort configuration remains a separate result; do not collapse these rows into a neutral per-model winner. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.
CursorBench 3.2 configurations are published exactly as Cursor reports them, including cost, token, and step averages. No per-model maximum is selected. Grok 4.5 rows carry Cursor's training-data-advantage disclosure and are not neutral winner evidence.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
37.6
Other version snapshots
Every benchmark version stays in a separate dated table. Scores, ranks, and changes are not compared across versions.
CursorBench 3.1
Observed Jun 30, 2026 · 12 configurations.
#Published configuration and provenanceRelative scoreScore
Configuration: Kimi 2.5 (single reported configuration)
Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource
Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.
31.9
How to read this benchmark
This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.
Trust this score when
The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
Cost, token, and step averages are evaluated together with the score for the same configuration.
You validate the configuration in your own Cursor workload before making a model choice.
Be cautious when
Rows are collapsed into one score per base model or compared across benchmark versions.
Small score differences are treated as decisive despite Cursor's variance caveat.
A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.
FAQ
What does the CursorBench benchmark measure?
Cursor's proprietary coding-agent benchmark for evaluating IDE-native multi-file coding workflows. LLMReference publishes CursorBench 3.2 at the configuration level with score, cost, token, step, and qualification provenance while retaining CursorBench 3.1 as dated history. Scores are useful for Cursor product context but are vendor-reported and not independently reproducible from a public harness. On this page it lists 42 tracked configurations where higher is better.
Is a higher CursorBench score always better?
For this benchmark, higher is better, but Cursor cautions that small differences may not be statistically meaningful. Treat each row as one exact configuration, not a neutral per-model winner, and validate qualified results such as Grok 4.5 separately.
How current is this CursorBench data?
This benchmark was last reviewed on Jul 18, 2026. Re-check the linked model pages for the freshest provider and pricing detail.