SWE-bench Pro
731-task multilingual real-world GitHub issue benchmark extending SWE-bench Verified with harder, more diverse tasks across Python, JavaScript, TypeScript, Java, Go, C++, and Rust.
Models ranked
46
tracked on this benchmark
Score band
80.3 – 38.7
best → lowest tracked
Snapshot trend
-5.20
Aug 12 → Aug 26 · 1 models
Leaderboard
Tracked models ranked by % Resolved (higher is better).
Notes: DAT-5842: Unstarred in Anthropic's shared Fable 5/Mythos 5 launch table, so this is a Claude Fable 5 score.
Configuration: Claude Opus 5 standard configuration
Notes: Vendor-reported. Source: p.149, Table 8.1.A. This is distinct from SWE-bench Verified.
Notes: Official Anthropic launch benchmark.
Notes: xAI launch blog SWE Bench Pro chart reports Grok 4.5 at 64.7% resolve rate. Evaluator/source: xAI first-party launch chart only; blog does not specify Pass@1 vs other metric, agent scaffold, reasoning level, or public SWE-bench Pro harness for Grok 4.5 (separate token-efficiency section references SWE Bench Pro tasks but not this score's run config). Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 64.7 pending harness documentation or independent leaderboard confirmation.
Notes: Official GPT-5.6 GA launch benchmark row.
Notes: Official LongCat-2.0 launch chart reports 59.51. Evaluator/harness details are not fully reproducible from public text; keep separate from independent SWE-bench Verified. Confidence: confirmed for reported value, medium for reproducibility.
Notes: DAT-5798: Self-reported in model card and SiliconFlow blog. Not independently verified.
Notes: Medium confidence secondary-source row for GPT-5.5 standard; attributed to Pro because the Pro product uses the same underlying weights.
Notes: OpenAI GPT-5.3-Codex appendix reports GPT-5.2-Codex 56.4%, GPT-5.2 55.6%, GPT-5.1-Codex-Max 50.8%. Evaluator: OpenAI self-reported launch benchmark. Harness details are not fully reproducible from public text; keep distinct from independent SWE-bench Verified. Confidence: confirmed for reported value, medium for reproducibility.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Related benchmarks
Last reviewed: Apr 15, 2026