DeepSWE 1.1
DeepSWE 1.1 is Datacurve's June 2026 revised execution and grading setup for the same long-horizon DeepSWE tasks, grading committed patches in a clean isolated verifier environment and using mini-swe-agent for the public leaderboard. Model scores must preserve the mini-swe-agent harness/evaluator context.
Models ranked
16
tracked on this benchmark
Score band
74.2 – 42.2
best → lowest tracked
Snapshot trend
+2.55
Aug 26 → Sep 22 · 5 models
Leaderboard
Tracked models ranked by Pass@1 (higher is better).
Configuration: gpt-6-astra (Extra High)
Notes: Datacurve DeepSWE 1.1 leaderboard (113 tasks, mini-swe-agent, Pass@1 over 4 runs). Verified on https://deepswe.datacurve.ai/ (page generated_at 2026-09-22; page shows integer %%). Precise value from Epoch AI mirror deepswe_external.csv (Source column = deepswe.datacurve.ai). Highest-scoring published effort selected.
Configuration: gemini-3.8-flash (High)
Notes: Datacurve DeepSWE 1.1 leaderboard (113 tasks, mini-swe-agent, Pass@1 over 4 runs). Verified on https://deepswe.datacurve.ai/ (page generated_at 2026-09-22; page shows integer %%). Precise value from Epoch AI mirror deepswe_external.csv (Source column = deepswe.datacurve.ai). Highest-scoring published effort selected.
Notes: Official GPT-5.6 GA launch benchmark row.
Configuration: Maximum effort
Notes: Vendor-reported. Source: pp.149-150, Table 8.1.A and Figure 8.2. The effort sweep ranges from 57.7 at low to 68.8 at maximum.
Configuration: kimi-k3 (Max)
Notes: Datacurve DeepSWE 1.1 leaderboard (113 tasks, mini-swe-agent, Pass@1 over 4 runs). Verified on https://deepswe.datacurve.ai/ (page generated_at 2026-09-22; page shows integer %%). Precise value from Epoch AI mirror deepswe_external.csv (Source column = deepswe.datacurve.ai). Highest-scoring published effort selected.
Notes: xAI launch post evals table reports Grok 4.6 High at 65.9% on DeepSWE v1.1. Variant: Grok 4.6 High (same model ID grok-4.6; high is the API default reasoning_effort).
Configuration: high thinking
Configuration: muse-spark-1.2 (Extra High)
Notes: Datacurve DeepSWE 1.1 leaderboard (113 tasks, mini-swe-agent, Pass@1 over 4 runs). Verified on https://deepswe.datacurve.ai/ (page generated_at 2026-09-22; page shows integer %%). Precise value from Epoch AI mirror deepswe_external.csv (Source column = deepswe.datacurve.ai). Highest-scoring published effort selected.
Notes: xAI launch blog DeepSWE 1.1 chart reports Grok 4.5 at 53% Pass@1. Evaluator/source: Datacurve with mini-swe-agent harness per xAI first-party claim. Variant: Grok 4.5 base API model. Harness status: mini-swe-agent (DeepSWE 1.1 public leaderboard setup). Confidence: medium for vendor-reported chart value, medium for harness identifiability. Recommended seed value: 53.
Configuration: gemini-3.6-flash (High)
Notes: Datacurve DeepSWE 1.1 leaderboard (113 tasks, mini-swe-agent, Pass@1 over 4 runs). Verified on https://deepswe.datacurve.ai/ (page generated_at 2026-09-22; page shows integer %%). Precise value from Epoch AI mirror deepswe_external.csv (Source column = deepswe.datacurve.ai). Highest-scoring published effort selected.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Related benchmarks
Last reviewed: Jul 8, 2026