activeCodingAgents

CursorBench

Metric: Score (higher is better)Introduced: 2026

This page is LLMReference's October 5, 2026 historical snapshot of CursorBench 4.0, not the current CursorBench release. Use it to compare Cursor IDE-native coding-agent configurations for multi-file work when Cursor is the target Coding or Agents workflow; it does not establish a neutral base-model winner. Read each score with that configuration's cost, token, step, and qualification context.

Configurations

46

tracked in CursorBench 4.0

Score band

57.8 – 16.0

best → lowest tracked

Snapshot trend

+17.50

CursorBench 4.0 · Sep 10 → Oct 5 · 3 models

Published snapshot

CursorBench 4.0 · observed Oct 5, 2026 · 46 configurations.

CursorBench 3.2 configurations are published exactly as Cursor reports them, including cost, token, and step averages. No per-model maximum is selected. Grok 4.5 rows carry Cursor's training-data-advantage disclosure and are not neutral winner evidence.

Compare candidates
#Published configuration and provenanceScore
1
Claude Opus 5.5

Configuration: claude-opus-5-5 (Max)

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Oct 5, 2026Confidence: confirmedCost/task: $13.43Tokens/task: 218,363Steps/task: 185Source

Notes: Cursor vendor-reported CursorBench 4.0 configuration result; verified on https://cursor.com/cursorbench chart label 'Opus 5.5 Max' (retrieved 2026-10-05). Not independently reproducible. Highest-scoring published configuration selected; timestamp is retrieval date (page carries no per-row date). Cost/tokens/steps from cursor.com/cursorbench table row (same retrieval).

57.8
2
Claude Sonnet 5.5

Configuration: claude-sonnet-5-5 (Max)

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Oct 5, 2026Confidence: confirmedCost/task: $9.67Tokens/task: 271,920Steps/task: 170Source

Notes: Cursor vendor-reported CursorBench 4.0 configuration result; verified on https://cursor.com/cursorbench chart label 'Sonnet 5.5 Max' (retrieved 2026-10-05). Not independently reproducible. Highest-scoring published configuration selected; timestamp is retrieval date (page carries no per-row date). Cost/tokens/steps from cursor.com/cursorbench table row (same retrieval).

55.5
3
Claude Fable 5.1

Configuration: Fable 5.1 Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $17.28Tokens/task: 117,236Steps/task: 128Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

51.8
4
Claude Fable 5.1

Configuration: Fable 5.1 Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $13.01Tokens/task: 87,294Steps/task: 101Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

51.6
5
Claude Fable 5.1

Configuration: Fable 5.1 High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $9.08Tokens/task: 58,438Steps/task: 77Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

49.2
6
Claude Fable 5.1

Configuration: Fable 5.1 Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $7.05Tokens/task: 45,411Steps/task: 63Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.8
7
Claude Opus 5

Configuration: Opus 5 Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $11.95Tokens/task: 85,384Steps/task: 106Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.6
8
Grok 4.7

Configuration: grok-4.7 (Extra High)

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Oct 5, 2026Confidence: confirmedCost/task: $6.01Tokens/task: 70,141Steps/task: 88Source

Notes: Cursor vendor-reported CursorBench 4.0 configuration result; verified on https://cursor.com/cursorbench chart label 'Grok 4.7 Extra High' (retrieved 2026-10-05). Not independently reproducible. Highest-scoring published configuration selected; timestamp is retrieval date (page carries no per-row date). Cost/tokens/steps from cursor.com/cursorbench table row (same retrieval).

46.3
9
Claude Opus 5

Configuration: Opus 5 Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $11.43Tokens/task: 80,094Steps/task: 103Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.1
10
Claude Fable 5.1

Configuration: Fable 5.1 Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $5.44Tokens/task: 34,795Steps/task: 51Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

45.1
11
Claude Opus 5

Configuration: Opus 5 High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $9.00Tokens/task: 61,405Steps/task: 86Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

44.7
12
Claude Opus 5

Configuration: Opus 5 Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $6.94Tokens/task: 45,272Steps/task: 72Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

43.3
13
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $8.23Tokens/task: 42,944Steps/task: 99Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

41.7
14
Muse Spark 1.3

Configuration: Muse Spark 1.3 Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $2.64Tokens/task: 52,005Steps/task: 98Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

41.6
15
Grok 4.6

Configuration: Grok 4.6 Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $6.10Tokens/task: 49,814Steps/task: 56Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

41.4
16
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $5.14Tokens/task: 60,814Steps/task: 107Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

41.3
17
Claude Opus 5

Configuration: Opus 5 Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $4.87Tokens/task: 31,995Steps/task: 57Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

40.7
18
Grok 4.6

Configuration: Grok 4.6 High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $5.20Tokens/task: 41,387Steps/task: 48Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

40.4
19
Gemini 3.8 Flash

Configuration: Gemini 3.8 Flash High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $4.70Tokens/task: 162,565Steps/task: 324Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

39.6
20
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $4.40Tokens/task: 24,729Steps/task: 55Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.7
21
Muse Spark 1.3

Configuration: Muse Spark 1.3 Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $2.10Tokens/task: 40,891Steps/task: 83Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.5
22
Gemini 3.8 Flash

Configuration: Gemini 3.8 Flash Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $4.06Tokens/task: 128,364Steps/task: 290Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.3
23
Grok 4.6

Configuration: Grok 4.6 Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $3.48Tokens/task: 24,893Steps/task: 40Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

36.1
24
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.03Tokens/task: 87,284Steps/task: 208Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

35.9
25
GPT-5.6 Sol

Configuration: GPT-5.6 Sol High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $2.85Tokens/task: 16,174Steps/task: 41Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

35.7
26
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $7.17Tokens/task: 149,257Steps/task: 140Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

34.1
27
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.81Tokens/task: 23,436Steps/task: 43Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

33.6
28
Grok 4.6

Configuration: Grok 4.6 Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $2.25Tokens/task: 16,307Steps/task: 32Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

33.4
29
Muse Spark 1.3

Configuration: Muse Spark 1.3 High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.66Tokens/task: 30,654Steps/task: 69Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

33.4
30
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.44Tokens/task: 40,598Steps/task: 98Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

33.0
31
Muse Spark 1.3

Configuration: Muse Spark 1.3 Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.49Tokens/task: 27,255Steps/task: 64Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

32.6
32
Claude Sonnet 5

Configuration: Sonnet 5 Extra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $4.55Tokens/task: 83,373Steps/task: 102Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

32.0
33
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.77Tokens/task: 10,111Steps/task: 32Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

31.1
34
Claude Sonnet 5

Configuration: Sonnet 5 High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $3.48Tokens/task: 61,146Steps/task: 85Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

30.8
35
GPT-5.6 Terra

Configuration: GPT-5.6 Terra High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.11Tokens/task: 13,162Steps/task: 33Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

30.7
36
GPT-5.6 Luna

Configuration: GPT-5.6 Luna High

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.25Tokens/task: 23,368Steps/task: 64Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

29.4
37
Muse Spark 1.3

Configuration: Muse Spark 1.3 Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.93Tokens/task: 17,483Steps/task: 47Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

29.3
38
Claude Sonnet 5

Configuration: Sonnet 5 Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $2.31Tokens/task: 39,114Steps/task: 65Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

28.0
39
Composer 2.5

Configuration: Composer 2.5

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.68Tokens/task: 17,347Steps/task: 41Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

27.7
40
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.64Tokens/task: 7,307Steps/task: 25Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

27.6
41
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.52Tokens/task: 5,914Steps/task: 23Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

25.2
42
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.87Tokens/task: 4,885Steps/task: 21Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

24.6
43
Muse Spark 1.3

Configuration: Muse Spark 1.3 Minimal

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.56Tokens/task: 10,620Steps/task: 34Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

24.3
44
Claude Sonnet 5

Configuration: Sonnet 5 Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $1.39Tokens/task: 23,772Steps/task: 46Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

24.1
45
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Medium

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.08Tokens/task: 7,642Steps/task: 32Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

22.2
46
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Low

Version: CursorBench 4.0Harness: CursorBench 4.0 productized Cursor-agent workflowEvaluator: CursorObserved: Sep 10, 2026Confidence: confirmedCost/task: $0.03Tokens/task: 3,288Steps/task: 18Source

Notes: Cursor vendor-reported configuration result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

16.0

Other version snapshots

Every benchmark version stays in a separate dated table. Scores, ranks, and changes are not compared across versions.

CursorBench 3.2, Grok 4.6 High (xAI first-party evals table)

Observed Aug 12, 2026 · 1 configuration.

#Published configuration and provenanceScore
1
Grok 4.6

Configuration: Grok 4.6 High

Version: CursorBench 3.2, Grok 4.6 High (xAI first-party evals table)Harness: Not recordedEvaluator: Not recordedObserved: Aug 12, 2026Confidence: Not recordedSource

Notes: xAI launch post evals table reports Grok 4.6 High at 69.9% on CursorBench v3.2.

69.9

CursorBench 3.2

Observed Jul 18, 2026 · 42 configurations.

#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $17.32Tokens/task: 103,525Steps/task: 72Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

70.5
2
Claude Fable 5

Configuration: Fable 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $11.73Tokens/task: 64,971Steps/task: 56Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

68.4
3
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.69Tokens/task: 28,320Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

67.2
4
Grok 4.5

Configuration: Grok 4.5 High*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 19,521Steps/task: 33Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.7
5
Claude Fable 5

Configuration: Fable 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $8.77Tokens/task: 43,747Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.5
6
Grok 4.5

Configuration: Grok 4.5 Medium*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.54Tokens/task: 18,914Steps/task: 34Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.4
7
Claude Fable 5

Configuration: Fable 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.80Tokens/task: 30,366Steps/task: 41Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.2
8
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.89Tokens/task: 32,969Steps/task: 47Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.9
9
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.88Tokens/task: 19,699Steps/task: 38Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.5
10
Grok 4.5

Configuration: Grok 4.5 Low*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.22Tokens/task: 15,841Steps/task: 31Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
11
GPT-5.6 Sol

Configuration: GPT-5.6 Sol High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.79Tokens/task: 13,867Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
12
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.77Tokens/task: 71,411Steps/task: 44Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.3
13
Claude Fable 5

Configuration: Fable 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.46Tokens/task: 18,182Steps/task: 31Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.1
14
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.45Tokens/task: 92,882Steps/task: 86Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.5
15
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.97Tokens/task: 87,973Steps/task: 61Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.1
16
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.95Tokens/task: 9,747Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

60.0
17
Claude Opus 4.8

Configuration: Opus 4.8 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.50Tokens/task: 51,121Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.4
18
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.44Tokens/task: 16,089Steps/task: 29Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.2
19
Claude Sonnet 5

Configuration: Sonnet 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.16Tokens/task: 52,871Steps/task: 67Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.7
20
GPT-5.5

Configuration: GPT-5.5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.05Tokens/task: 12,183Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
21
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.85Tokens/task: 17,534Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
22
Claude Opus 4.8

Configuration: Opus 4.8 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.15Tokens/task: 33,548Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.0
23
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.14Tokens/task: 22,480Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

57.7
24
Claude Sonnet 5

Configuration: Sonnet 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.19Tokens/task: 39,483Steps/task: 57Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.9
25
GPT-5.6 Luna

Configuration: GPT-5.6 Luna High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.82Tokens/task: 15,141Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.8
26
Claude Opus 4.8

Configuration: Opus 4.8 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.81Tokens/task: 28,384Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
27
Composer 2.5

Configuration: Composer 2.5

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.44Tokens/task: 14,286Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
28
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.76Tokens/task: 35,946Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

55.0
29
GPT-5.6 Terra

Configuration: GPT-5.6 Terra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.89Tokens/task: 9,468Steps/task: 23Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

54.2
30
GPT-5.5

Configuration: GPT-5.5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 8,522Steps/task: 25Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.8
31
Claude Opus 4.8

Configuration: Opus 4.8 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.02Tokens/task: 19,624Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.1
32
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.01Tokens/task: 5,104Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.6
33
Claude Sonnet 5

Configuration: Sonnet 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.16Tokens/task: 26,200Steps/task: 46Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.4
34
GLM-5.2

Configuration: GLM 5.2 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.19Tokens/task: 21,829Steps/task: 49Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

51.5
35
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.61Tokens/task: 6,222Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

50.3
36
Kimi K2.7-Code

Configuration: Kimi K2.7 Code

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.43Tokens/task: 31,247Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

49.7
37
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.20Tokens/task: 46,702Steps/task: 77Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

48.8
38
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.39Tokens/task: 7,095Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
39
Claude Sonnet 5

Configuration: Sonnet 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.30Tokens/task: 16,269Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
40
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.53Tokens/task: 5,312Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.9
41
GPT-5.5

Configuration: GPT-5.5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.98Tokens/task: 5,168Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.6
42
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.16Tokens/task: 3,209Steps/task: 17Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.6

CursorBench 3.1

Observed Jun 30, 2026 · 12 configurations.

#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

72.9
2
Claude Opus 4.7

Configuration: Opus 4.7 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.8
3
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.3
4
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

63.8
5
Composer 2.5

Configuration: Composer 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

63.2
6
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

61.2
7
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

54.6
8
Composer 2

Configuration: Composer 2 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

52.2
9
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

49.8
10
Claude Sonnet 4.6

Configuration: Sonnet 4.6 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

49.0
11
Kimi K2.6

Configuration: Kimi 2.6 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

47.6
12
Kimi K2.5

Configuration: Kimi 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

31.9

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
  • Cost, token, and step averages are evaluated together with the score for the same configuration.
  • You validate the configuration in your own Cursor workload before making a model choice.

Be cautious when

  • Rows are collapsed into one score per base model or compared across benchmark versions.
  • Small score differences are treated as decisive despite Cursor's variance caveat.
  • A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.

Related benchmarks

Last reviewed: Sep 14, 2026

Resources