LLM Reference
activeCodingAgents

CursorBench

Metric: Score (higher is better)Introduced: 2026

The 42 rows below are LLMReference's July 18, 2026 snapshot of CursorBench 3.2. CursorBench is a vendor-run evaluation of IDE-native, multi-file coding-agent workflows, not standalone base models, and its public harness is not independently reproducible. Each reasoning or effort configuration remains a separate result; do not collapse these rows into a neutral per-model winner.

Configurations

42

tracked in CursorBench 3.2

Score band

70.5 – 37.6

best → lowest tracked

Snapshot trend

CursorBench 3.2 · need ≥2 snapshots

Published snapshot

CursorBench 3.2 · observed Jul 18, 2026 · 42 configurations.

CursorBench 3.2 configurations are published exactly as Cursor reports them, including cost, token, and step averages. No per-model maximum is selected. Grok 4.5 rows carry Cursor's training-data-advantage disclosure and are not neutral winner evidence.

Compare candidates
#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $17.32Tokens/task: 103,525Steps/task: 72Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

70.5
2
Claude Fable 5

Configuration: Fable 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $11.73Tokens/task: 64,971Steps/task: 56Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

68.4
3
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.69Tokens/task: 28,320Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

67.2
4
Grok 4.5

Configuration: Grok 4.5 High*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 19,521Steps/task: 33Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.7
5
Claude Fable 5

Configuration: Fable 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $8.77Tokens/task: 43,747Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

66.5
6
Grok 4.5

Configuration: Grok 4.5 Medium*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.54Tokens/task: 18,914Steps/task: 34Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.4
7
Claude Fable 5

Configuration: Fable 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.80Tokens/task: 30,366Steps/task: 41Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

65.2
8
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.89Tokens/task: 32,969Steps/task: 47Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.9
9
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.88Tokens/task: 19,699Steps/task: 38Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

64.5
10
Grok 4.5

Configuration: Grok 4.5 Low*

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.22Tokens/task: 15,841Steps/task: 31Source

Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
11
GPT-5.6 Sol

Configuration: GPT-5.6 Sol High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.79Tokens/task: 13,867Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

63.5
12
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $5.77Tokens/task: 71,411Steps/task: 44Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.3
13
Claude Fable 5

Configuration: Fable 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.46Tokens/task: 18,182Steps/task: 31Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

62.1
14
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $6.45Tokens/task: 92,882Steps/task: 86Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.5
15
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.97Tokens/task: 87,973Steps/task: 61Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

61.1
16
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.95Tokens/task: 9,747Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

60.0
17
Claude Opus 4.8

Configuration: Opus 4.8 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.50Tokens/task: 51,121Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.4
18
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.44Tokens/task: 16,089Steps/task: 29Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

59.2
19
Claude Sonnet 5

Configuration: Sonnet 5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $4.16Tokens/task: 52,871Steps/task: 67Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.7
20
GPT-5.5

Configuration: GPT-5.5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.05Tokens/task: 12,183Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
21
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.85Tokens/task: 17,534Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.4
22
Claude Opus 4.8

Configuration: Opus 4.8 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.15Tokens/task: 33,548Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

58.0
23
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Extra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.14Tokens/task: 22,480Steps/task: 48Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

57.7
24
Claude Sonnet 5

Configuration: Sonnet 5 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $3.19Tokens/task: 39,483Steps/task: 57Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.9
25
GPT-5.6 Luna

Configuration: GPT-5.6 Luna High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.82Tokens/task: 15,141Steps/task: 40Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.8
26
Claude Opus 4.8

Configuration: Opus 4.8 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.81Tokens/task: 28,384Steps/task: 32Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
27
Composer 2.5

Configuration: Composer 2.5

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.44Tokens/task: 14,286Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

56.1
28
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.76Tokens/task: 35,946Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

55.0
29
GPT-5.6 Terra

Configuration: GPT-5.6 Terra High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.89Tokens/task: 9,468Steps/task: 23Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

54.2
30
GPT-5.5

Configuration: GPT-5.5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.51Tokens/task: 8,522Steps/task: 25Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.8
31
Claude Opus 4.8

Configuration: Opus 4.8 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.02Tokens/task: 19,624Steps/task: 27Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

53.1
32
GPT-5.6 Sol

Configuration: GPT-5.6 Sol Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.01Tokens/task: 5,104Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.6
33
Claude Sonnet 5

Configuration: Sonnet 5 Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.16Tokens/task: 26,200Steps/task: 46Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

52.4
34
GLM-5.2

Configuration: GLM 5.2 High

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.19Tokens/task: 21,829Steps/task: 49Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

51.5
35
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.61Tokens/task: 6,222Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

50.3
36
Kimi K2.7-Code

Configuration: Kimi K2.7 Code

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.43Tokens/task: 31,247Steps/task: 58Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

49.7
37
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $2.20Tokens/task: 46,702Steps/task: 77Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

48.8
38
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Medium

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.39Tokens/task: 7,095Steps/task: 28Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
39
Claude Sonnet 5

Configuration: Sonnet 5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $1.30Tokens/task: 16,269Steps/task: 33Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

47.7
40
GPT-5.6 Terra

Configuration: GPT-5.6 Terra Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.53Tokens/task: 5,312Steps/task: 19Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.9
41
GPT-5.5

Configuration: GPT-5.5 Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.98Tokens/task: 5,168Steps/task: 20Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

46.6
42
GPT-5.6 Luna

Configuration: GPT-5.6 Luna Low

Version: CursorBench 3.2Harness: CursorBench 3.2 productized Cursor-agent workflowEvaluator: CursorObserved: Jul 18, 2026Confidence: confirmedCost/task: $0.16Tokens/task: 3,209Steps/task: 17Source

Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.

37.6

Other version snapshots

Every benchmark version stays in a separate dated table. Scores, ranks, and changes are not compared across versions.

CursorBench 3.2, Grok 4.6 High (xAI first-party evals table)

Observed Aug 12, 2026 · 1 configuration.

#Published configuration and provenanceScore
1
Grok 4.6

Configuration: Grok 4.6 High

Version: CursorBench 3.2, Grok 4.6 High (xAI first-party evals table)Harness: Not recordedEvaluator: Not recordedObserved: Aug 12, 2026Confidence: Not recordedSource

Notes: xAI launch post evals table reports Grok 4.6 High at 69.9% on CursorBench v3.2.

69.9

CursorBench 3.1

Observed Jun 30, 2026 · 12 configurations.

#Published configuration and provenanceScore
1
Claude Fable 5

Configuration: Fable 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

72.9
2
Claude Opus 4.7

Configuration: Opus 4.7 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.8
3
GPT-5.5

Configuration: GPT-5.5 Extra High

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

64.3
4
Claude Opus 4.8

Configuration: Opus 4.8 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

63.8
5
Composer 2.5

Configuration: Composer 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

63.2
6
Claude Sonnet 5

Configuration: Sonnet 5 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

61.2
7
GLM-5.2

Configuration: GLM 5.2 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

54.6
8
Composer 2

Configuration: Composer 2 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

52.2
9
Gemini 3.5 Flash

Configuration: Gemini 3.5 Flash (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

49.8
10
Claude Sonnet 4.6

Configuration: Sonnet 4.6 Max

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.

49.0
11
Kimi K2.6

Configuration: Kimi 2.6 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

47.6
12
Kimi K2.5

Configuration: Kimi 2.5 (single reported configuration)

Version: CursorBench 3.1Harness: CursorBench 3.1Evaluator: CursorObserved: Jun 30, 2026Confidence: confirmedSource

Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.

31.9

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • The row names the exact model configuration, benchmark version, evaluator, harness, and observation date.
  • Cost, token, and step averages are evaluated together with the score for the same configuration.
  • You validate the configuration in your own Cursor workload before making a model choice.

Be cautious when

  • Rows are collapsed into one score per base model or compared across benchmark versions.
  • Small score differences are treated as decisive despite Cursor's variance caveat.
  • A qualified row, including Grok 4.5's disclosed training-data advantage, is used as neutral winner evidence.

Related benchmarks

Last reviewed: Jul 18, 2026

Resources