LLM Reference

Claude Opus 5

Released
2026-07-24
Last refreshed
2026-07-26
Status
Researched 45d ago
ProprietaryCommercial use: conditionalMultimodalCodingRAGAgentsLong contextVisionJSON / Tool useHighlight

Claude Opus 5 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 1m context window
  • Buyers comparing 4 tracked provider routes

Do not use it for

  • Workloads where another current model has stronger sourced task evidence
Specifications
Family
Claude 5
Released
2026-07-24
Context
1m
Max output
128,000
Architecture
Decoder Only
Knowledge cutoff
2026-05
Specialization
general
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Not released
Training
Pretrained
Created by

Developing safe and ethical AI systems.

San Francisco, California, United States
Founded 2021
Website
Pricing
Output / 1M
$25.00
Input / 1M
$5.00

Cheapest of 7 routes · Anthropic · cache read $0.500

About

Claude Opus 5 is Anthropic's July 2026 flagship Opus model for complex agentic coding, enterprise work, long-horizon reasoning, computer use, and professional analysis. It accepts text and image input and returns text, with adaptive thinking enabled by default and five request-level effort settings from low through max that control thinking depth and token use rather than visible response length. It also supports tool use, structured outputs, prompt caching, Batch API, a 1M-token context window, and up to 128K synchronous output tokens.

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ D

2 relevant benchmarks in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

Q/$ D

1 relevant benchmark in the decision map.

Provider price ladder

Compare all 7

Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MBatch in / outCacheRoute
Anthropic$5.00$25.00$2.50 / $12.50read $0.500 / 5m $6.25 / 1h $10.00
Serverless
GCP Vertex AI$5.00$25.00$2.50 / $12.50read $0.500 / 5m $6.25 / 1h $10.00
Serverless
OpenRouter$5.00$25.00--
Serverless
Vercel AI Gateway$5.00$25.00-read $0.500 / 5m $6.25
Serverless

Available via routers & gateways(16)

Capabilities

VisionMultimodalReasoningJSON / Tool useStructured OutputsCode ExecutionPrompt CachingBatch API

Benchmark peer barsfor Coding

Benchmark scores(10)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
SWE-bench Verified96.0Verified, 500-problem solvable subsetObserved 2026-07-24
Configuration: Adaptive thinking at maximum effort
Harness: Anthropic system-card standard configuration; average of five trials
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: p.149, Table 8.1.A. Preserve the five-trial setting; do not treat single-trial rows as equivalent.
Source
SWE-bench Pro79.2SWE-bench ProObserved 2026-07-24
Configuration: Claude Opus 5 standard configuration
Harness: Anthropic system-card standard configuration; average of five trials
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: p.149, Table 8.1.A. This is distinct from SWE-bench Verified.
Source
SWE-bench Multilingual89.5SWE-bench MultilingualObserved 2026-07-24
Configuration: Claude Opus 5 standard configuration
Harness: Anthropic system-card standard configuration; average of five trials
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: p.149, Table 8.1.A. This multilingual suite is not the English SWE-bench Verified row.
Source
SWE-bench Multimodal59.4SWE-bench MultimodalObserved 2026-07-24
Configuration: Claude Opus 5 standard configuration
Harness: Anthropic internal SWE-bench Multimodal harness
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: p.149, Table 8.1.A and Section 9.3. Do not compare this internal-harness row as identical to third-party runs.
Source
DeepSWE 1.168.8DeepSWE v1.1Observed 2026-07-24
Configuration: Maximum effort
Harness: 113 long-horizon agentic software-engineering tasks; five-trial average
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: pp.149-150, Table 8.1.A and Figure 8.2. The effort sweep ranges from 57.7 at low to 68.8 at maximum.
Source
FrontierCode 1.1 Main53.4FrontierCode 1.1 MainObserved 2026-07-24
Configuration: Medium effort
Harness: 100-task mean@5 composite of held-out-test performance and weighted code quality
Evaluator: Cognition
Confidence: confirmed
Notes: Source: p.150, Figure 8.3. Medium is the reported Opus 5 setting; preserve it and keep Main separate from Extended.
Source
FrontierCode 1.1 Extended63.6FrontierCode 1.1 ExtendedObserved 2026-07-24
Configuration: Medium effort
Harness: Cognition FrontierCode 1.1 weighted code-quality and held-out-test mean@5 evaluation
Evaluator: Cognition
Confidence: confirmed
Notes: Source: p.151, Figure 8.4. Keep Extended separate from FrontierCode Main.
Source
Frontier-Bench v0.144.4Frontier-Bench v0.1Observed 2026-07-24
Configuration: xhigh effort
Harness: 74 tasks; mini-SWE-agent on Google Kubernetes Engine; mean reward over five attempts per task
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: p.152, Figure 8.5. About 5% of API calls were refused and 4% of trials fell back to Opus 4.8; maximum effort scored 43.0 within noise of xhigh.
Source
ArXivMath June 202690.8June 2026Observed 2026-07-24
Configuration: Maximum effort, no tools
Harness: 49-problem June 2026 set; four runs per problem
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 91.33 with-tools configuration.
Source
ArXivMath June 202691.3June 2026Observed 2026-07-24
Configuration: Maximum effort, with tools
Harness: 49-problem June 2026 set; four runs per problem
Evaluator: Anthropic
Confidence: confirmed
Notes: Vendor-reported. Source: pp.154-155, Figure 8.10. Store separately from the 90.82 no-tools configuration.
Source

Migration checks

No linked migration route is available for this model yet.

API versions

claude-opus-5