LLM Reference
activeCodingAgents

CursorBench

Metric: Score (higher is better)Introduced: 2026

The 42 rows below are LLMReference's July 18, 2026 snapshot of CursorBench 3.2. CursorBench is a vendor-run evaluation of IDE-native, multi-file coding-agent workflows, not standalone base models, and its public harness is not independently reproducible. Each reasoning or effort configuration remains a separate result; do not collapse these rows into a neutral per-model winner. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Configurations

42

tracked in CursorBench 3.2

Score band

70.5 – 37.6

best → lowest tracked

Snapshot trend

CursorBench 3.2 · need ≥2 snapshots

Published snapshot

CursorBench 3.2 · observed Jul 18, 2026 · 42 configurations.

CursorBench 3.2 configurations are published exactly as Cursor reports them, including cost, token, and step averages. No per-model maximum is selected. Grok 4.5 rows carry Cursor's training-data-advantage disclosure and are not neutral winner evidence.

Compare candidates
#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $17.32Tokens/task: 103,525Steps/task: 72Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

70.5
2
Claude Fable 5

Configuration: Fable 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $11.73Tokens/task: 64,971Steps/task: 56Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

68.4
3
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.69Tokens/task: 28,320Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

67.2
4
Grok 4.5

Configuration: Grok 4.5 High*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 19,521Steps/task: 33Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.7
5
Claude Fable 5

Configuration: Fable 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $8.77Tokens/task: 43,747Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.5
6
Grok 4.5

Configuration: Grok 4.5 Medium*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.54Tokens/task: 18,914Steps/task: 34Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.4
7
Claude Fable 5

Configuration: Fable 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.80Tokens/task: 30,366Steps/task: 41Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.2
8
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.89Tokens/task: 32,969Steps/task: 47Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.9
9
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.88Tokens/task: 19,699Steps/task: 38Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.5
10
Grok 4.5

Configuration: Grok 4.5 Low*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.22Tokens/task: 15,841Steps/task: 31Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
11
GPT-5.6 Sol

Configuration: GPT-5.6 Sol High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.79Tokens/task: 13,867Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
12
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.77Tokens/task: 71,411Steps/task: 44Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.3
13
Claude Fable 5

Configuration: Fable 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.46Tokens/task: 18,182Steps/task: 31Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.1
14
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.45Tokens/task: 92,882Steps/task: 86Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.5
15
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.97Tokens/task: 87,973Steps/task: 61Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.1
16
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.95Tokens/task: 9,747Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

60.0
17
Claude Opus 4.8

Configuration: Opus 4.8 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.50Tokens/task: 51,121Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.4
18
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.44Tokens/task: 16,089Steps/task: 29Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.2
19
Claude Sonnet 5

Configuration: Sonnet 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.16Tokens/task: 52,871Steps/task: 67Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.7
20
GPT-5.5

Configuration: GPT-5.5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.05Tokens/task: 12,183Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
21
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.85Tokens/task: 17,534Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
22
Claude Opus 4.8

Configuration: Opus 4.8 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.15Tokens/task: 33,548Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.0
23
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.14Tokens/task: 22,480Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

57.7
24
Claude Sonnet 5

Configuration: Sonnet 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.19Tokens/task: 39,483Steps/task: 57Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.9
25
GPT-5.6 Luna

Configuration: GPT-5.6 Luna High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.82Tokens/task: 15,141Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.8
26
Claude Opus 4.8

Configuration: Opus 4.8 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.81Tokens/task: 28,384Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
27
Composer 2.5

Configuration: Composer 2.5

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.44Tokens/task: 14,286Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
28
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.76Tokens/task: 35,946Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

55.0
29
GPT-5.6 Terra

Configuration: GPT-5.6 Terra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.89Tokens/task: 9,468Steps/task: 23Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

54.2
30
GPT-5.5

Configuration: GPT-5.5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 8,522Steps/task: 25Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.8
31
Claude Opus 4.8

Configuration: Opus 4.8 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.02Tokens/task: 19,624Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.1
32
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.01Tokens/task: 5,104Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.6
33
Claude Sonnet 5

Configuration: Sonnet 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.16Tokens/task: 26,200Steps/task: 46Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.4
34
GLM-5.2

Configuration: GLM 5.2 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.19Tokens/task: 21,829Steps/task: 49Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

51.5
35
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.61Tokens/task: 6,222Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

50.3
36
Kimi K2.7-Code

Configuration: Kimi K2.7 Code

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.43Tokens/task: 31,247Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

49.7
37
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.20Tokens/task: 46,702Steps/task: 77Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

48.8
38
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.39Tokens/task: 7,095Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
39
Claude Sonnet 5

Configuration: Sonnet 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.30Tokens/task: 16,269Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
40
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.53Tokens/task: 5,312Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.9
41
GPT-5.5

Configuration: GPT-5.5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.98Tokens/task: 5,168Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.6
42
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.16Tokens/task: 3,209Steps/task: 17Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.6

Other version snapshots

Every benchmark version stays in a separate dated table. Scores, ranks, and changes are not compared across versions.

CursorBench 3.1

Observed Jun 30, 2026 · 12 configurations.

#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

72.9
2
Claude Opus 4.7

Configuration: Opus 4.7 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.8
3
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.3
4
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

63.8
5
Composer 2.5

Configuration: Composer 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

63.2
6
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

61.2
7
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

54.6
8
Composer 2

Configuration: Composer 2 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

52.2
9
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

49.8
10
Claude Sonnet 4.6

Configuration: Sonnet 4.6 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

49.0
11
Kimi K2.6

Configuration: Kimi 2.6 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

47.6
12
Kimi K2.5

Configuration: Kimi 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

31.9

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
  • Cost, token, and step averages are evaluated together with the score for the same configuration.
  • You validate the configuration in your own Cursor workload before making a model choice.

Be cautious when

  • Rows are collapsed into one score per base model or compared across benchmark versions.
  • Small score differences are treated as decisive despite Cursor's variance caveat.
  • A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.

FAQ

What does the CursorBench benchmark measure?

Cursor's proprietary coding-agent benchmark for evaluating IDE-native multi-file coding workflows. LLMReference publishes CursorBench 3.2 at the configuration level with score, cost, token, step, and qualification provenance while retaining CursorBench 3.1 as dated history. Scores are useful for Cursor product context but are vendor-reported and not independently reproducible from a public harness. On this page it lists 42 tracked configurations where higher is better.

Is a higher CursorBench score always better?

For this benchmark, higher is better, but Cursor cautions that small differences may not be statistically meaningful. Treat each row as one exact configuration, not a neutral per-model winner, and validate qualified results such as Grok 4.5 separately.

How current is this CursorBench data?

This benchmark was last reviewed on Jul 18, 2026. Re-check the linked model pages for the freshest provider and pricing detail.

Related benchmarks

Last reviewed: Jul 18, 2026

Resources