LLM Reference
activeCoding

SWE-bench Verified

Metric: % ResolvedIntroduced: 2024

500 human-validated GitHub issue resolution tasks from SWE-bench, created with OpenAI in August 2024. The standard evaluation for agentic coding systems. Top performers (2026) exceed 78% resolved. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Models ranked

81

tracked on this benchmark

Score band

96.0 – 23.6

best → lowest tracked

Snapshot trend

+10.80

Jun 30 → Jul 24 · 1 models

Leaderboard

Tracked models ranked by % Resolved (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Claude Fable 5
Version: Vals.ai harness (weighted avg by difficulty tier)Harness: Not recordedEvaluator: Not recordedObserved: Jun 9, 2026Confidence: Not recordedSource

Notes: DAT-5842: Independent Vals.ai SWE-bench Verified result; keep separate from Anthropic's SWE-bench Pro launch row.

96.0
2
Claude Opus 5

Configuration: Adaptive thinking at maximum effort

Version: Verified, 500-problem solvable subsetHarness: Anthropic system-card standard configuration; average of five trialsEvaluator: AnthropicObserved: Jul 24, 2026Confidence: confirmedSource

Notes: Vendor-reported. Source: p.149, Table 8.1.A. Preserve the five-trial setting; do not treat single-trial rows as equivalent.

96.0
3
Claude Mythos Preview
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: May 1, 2026Confidence: Not recordedSource
93.9
4
Claude Opus 4.8
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch benchmark.

88.6
5
Claude Opus 4.7
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
87.6
6
Claude Sonnet 5
Version: SWE-bench Verified; Anthropic standard configuration, adaptive thinking at max effort, default sampling, average over 5 trialsHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
85.2
7
GPT-5.3-Codex
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
85.0
8
GPT-5.5
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 30, 2026Confidence: Not recordedSource

Notes: DAT-5460 source reconciliation: vals.ai standardized harness reports GPT-5.5 at 82.6; do not use the non-comparable 88.7 self-reported row for /best/coding.

82.6
9
GPT-5.5 Pro
Version: Vals.ai independent SWE-bench Verified; standard GPT-5.5 weightsHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource

Notes: Medium confidence: GPT-5.5 Pro uses the same underlying weights as standard GPT-5.5; no Pro-specific SWE-bench Verified score was found.

82.6
10
Claude Opus 4.5
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Nov 1, 2025Confidence: Not recordedSource

Notes: Confidence: medium. DAT-4172 May 12 /best/ refresh; benchlm.ai May 11 snapshot showed Claude Opus 4.5 at 80.9%.

80.9
11
Claude Opus 4.6
Version: 2026-02Harness: Not recordedEvaluator: Not recordedObserved: Apr 9, 2026Confidence: Not recordedSource
80.8
12
DeepSeek V4 Pro
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
80.6
13
Gemini 3.1 Pro Preview
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
80.6
14
MiniMax M3
Version: Rank 8 of 49 (resolved%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
80.5
15
Qwen3.7-Max
Version: Rank 9 of 97 (resolved%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
80.4
16
Kimi K2.6
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
80.2
17
MiniMax M2.5 Highspeed
Version: M2 (resolved%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
80.2
18
GPT-5.2
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource

Notes: Verified via third-party source; official OpenAI benchmark may differ. Use directionally.

80.0
19
Claude Sonnet 4.6
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Feb 17, 2026Confidence: Not recordedSource

Notes: Confidence: high. DAT-4934 confirmed Anthropic self-reported this score and identified Composer 2.5's prior 80.8 SWE-bench Verified row as a Claude Opus 4.6 misattribution.

79.6
20
DeepSeek V4 Flash
Version: SWE-bench Verified resolved, V4-Flash Think MaxHarness: Not recordedEvaluator: Not recordedObserved: Jul 3, 2026Confidence: Not recordedSource
79.0
21
Xiaomi MiMo-V2.5-Pro
Version: SWE-Bench Verified (Resolved)Harness: Not recordedEvaluator: Not recordedObserved: Apr 28, 2026Confidence: Not recordedSource
78.9
22
Qwen3-Max
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: Apr 24, 2026Confidence: Not recordedSource
78.8
23
Qwen3.6-Plus
Version: SWE-bench VerifiedHarness: Not recordedEvaluator: Not recordedObserved: May 12, 2026Confidence: Not recordedSource

Notes: Confidence: high. DAT-4172 May 12 /best/ refresh; benchlm.ai and llm-stats both reported Qwen3.6 Plus at 78.8%.

78.8
24
Qwen3.6 Max Preview
Version: SWE-bench Verified (resolved)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
78.8
25
Gemini 3 Flash
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Dec 17, 2025Confidence: Not recordedSource
78.0
Showing the top 25 of 81 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

FAQ

What does the benchmark measure?

500 human-validated GitHub issue resolution tasks from SWE-bench, created with OpenAI in August 2024. The standard evaluation for agentic coding systems. Top performers (2026) exceed 78% resolved. On this page it lists 81 tracked model variants where higher is better.

Is a higher SWE-bench Verified score always better?

For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.

How current is this SWE-bench Verified data?

This benchmark was last reviewed on Apr 15, 2026. The tracked score average moved +10.80 points across the last 3 snapshots.

Related benchmarks

Last reviewed: Apr 15, 2026