CursorBench 3.2, Grok 4.6 High (xAI first-party evals table)
Observed Aug 12, 2026 · 1 configuration.
The 42 rows below are LLMReference's July 18, 2026 snapshot of CursorBench 3.2. CursorBench is a vendor-run evaluation of IDE-native, multi-file coding-agent workflows, not standalone base models, and its public harness is not independently reproducible. Each reasoning or effort configuration remains a separate result; do not collapse these rows into a neutral per-model winner.
Configurations
42
tracked in CursorBench 3.2
Score band
70.5 – 37.6
best → lowest tracked
Snapshot trend
—
CursorBench 3.2 · need ≥2 snapshots
CursorBench 3.2 · observed Jul 18, 2026 · 42 configurations.
CursorBench 3.2 configurations are published exactly as Cursor reports them, including cost, token, and step averages. No per-model maximum is selected. Grok 4.5 rows carry Cursor's training-data-advantage disclosure and are not neutral winner evidence.
Configuration: Fable 5 Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Fable 5 Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Sol Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Grok 4.5 High*
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Fable 5 High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Grok 4.5 Medium*
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Fable 5 Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Terra Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Sol Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Grok 4.5 Low*
Qualification: Cursor discloses a training-data advantage; do not use this result for a neutral ranking claim.
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Sol High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Opus 4.8 Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Fable 5 Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Sonnet 5 Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Luna Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Sol Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Opus 4.8 Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Terra Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Sonnet 5 Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.5 High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.5 Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Opus 4.8 High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Luna Extra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Sonnet 5 High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Luna High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Opus 4.8 Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Composer 2.5
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GLM 5.2 Max
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Terra High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.5 Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Opus 4.8 Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Sol Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Sonnet 5 Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GLM 5.2 High
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Terra Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Kimi K2.7 Code
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Gemini 3.5 Flash
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Luna Medium
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: Sonnet 5 Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Terra Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.5 Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Configuration: GPT-5.6 Luna Low
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Every benchmark version stays in a separate dated table. Scores, ranks, and changes are not compared across versions.
Observed Aug 12, 2026 · 1 configuration.
Observed Jun 30, 2026 · 12 configurations.
Configuration: Fable 5 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: Opus 4.7 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: GPT-5.5 Extra High
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: Opus 4.8 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: Composer 2.5 (single reported configuration)
Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.
Configuration: Sonnet 5 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: GLM 5.2 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Configuration: Composer 2 (single reported configuration)
Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.
Configuration: Gemini 3.5 Flash (single reported configuration)
Notes: Cursor published one CursorBench 3.1 configuration for this model; no cross-effort selection was needed.
Configuration: Sonnet 4.6 Max
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
Be cautious when
Last reviewed: Jul 18, 2026