SWE-bench Pro
731-task multilingual real-world GitHub issue benchmark extending SWE-bench Verified with harder, more diverse tasks across Python, JavaScript, TypeScript, Java, Go, C++, and Rust. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.
Models ranked
42
tracked on this benchmark
Score band
80.3 – 38.7
best → lowest tracked
Snapshot trend
+14.50
Jul 8 → Jul 24 · 1 models
Leaderboard
Tracked models ranked by % Resolved (higher is better).
Notes: DAT-5842: Unstarred in Anthropic's shared Fable 5/Mythos 5 launch table, so this is a Claude Fable 5 score.
Configuration: Claude Opus 5 standard configuration
Notes: Vendor-reported. Source: p.149, Table 8.1.A. This is distinct from SWE-bench Verified.
Notes: Official Anthropic launch benchmark.
Notes: xAI launch blog SWE Bench Pro chart reports Grok 4.5 at 64.7% resolve rate. Evaluator/source: xAI first-party launch chart only; blog does not specify Pass@1 vs other metric, agent scaffold, reasoning level, or public SWE-bench Pro harness for Grok 4.5 (separate token-efficiency section references SWE Bench Pro tasks but not this score's run config). Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 64.7 pending harness documentation or independent leaderboard confirmation.
Notes: Official GPT-5.6 GA launch benchmark row.
Notes: Official LongCat-2.0 launch chart reports 59.51. Evaluator/harness details are not fully reproducible from public text; keep separate from independent SWE-bench Verified. Confidence: confirmed for reported value, medium for reproducibility.
Notes: DAT-5798: Self-reported in model card and SiliconFlow blog. Not independently verified.
Notes: Medium confidence secondary-source row for GPT-5.5 standard; attributed to Pro because the Pro product uses the same underlying weights.
Notes: OpenAI GPT-5.3-Codex appendix reports GPT-5.2-Codex 56.4%, GPT-5.2 55.6%, GPT-5.1-Codex-Max 50.8%. Evaluator: OpenAI self-reported launch benchmark. Harness details are not fully reproducible from public text; keep distinct from independent SWE-bench Verified. Confidence: confirmed for reported value, medium for reproducibility.
How to read this benchmark
This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
FAQ
What does the benchmark measure?
731-task multilingual real-world GitHub issue benchmark extending SWE-bench Verified with harder, more diverse tasks across Python, JavaScript, TypeScript, Java, Go, C++, and Rust. On this page it lists 42 tracked model variants where higher is better.
Is a higher SWE-bench Pro score always better?
For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.
How current is this SWE-bench Pro data?
This benchmark was last reviewed on Apr 15, 2026. The tracked score average moved +14.50 points across the last 3 snapshots.
Related benchmarks
Last reviewed: Apr 15, 2026