Claude Opus 4.8
Claude Opus 4.8 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 1m context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- Claude 4.8
- Released
- 2026-05-28
- Context
- 1m
- Max output
- 128,000
- Architecture
- Decoder Only
- Knowledge cutoff
- 2026-01
- Specialization
- general
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Not released
- Training
- Fine-tuned
Cheapest of 6 routes · Anthropic · cache read $0.500
About
Claude Opus 4.8 is Anthropic's flagship Claude 4.8 model, released May 28, 2026 for agentic coding, long-horizon reasoning, computer use, and professional knowledge work. It supports text and image inputs, adaptive reasoning, tool use, structured outputs, computer-use tools, prompt caching, Batch API, Dynamic Workflows parallel subagents, a 1M-token context window on Anthropic API/Bedrock/Vertex, and 128K max output. Key datapack rows: SWE-bench Pro 69.2%, SWE-bench Verified 88.6%, Terminal-Bench 2.1 74.6%, HLE with tools 57.9%, OSWorld-Verified 83.4%, GDPval-AA 1890 Elo, and MCP-Atlas 82.2%. Standard Anthropic API pricing is $5/M input and $25/M output.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ D3 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ D1 relevant benchmark in the decision map.
Provider price ladder
Compare all 6Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Batch in / out | Cache | Route |
|---|---|---|---|---|---|
| Anthropic | $5.00 | $25.00 | $2.50 / $12.50 | read $0.500 / 5m $6.25 / 1h $10.00 | Serverless |
| OpenRouter | $5.00 | $25.00 | - | - | Serverless |
| AWS Bedrock | - | - | - | - | ServerlessPartial |
| GCP Vertex AI | - | - | - | - | ServerlessPartial |
Available via routers & gateways(16)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
AIRouter
RouterCommercial LLM router that analyzes incoming requests and routes to the optimal model for cost/quality/latency via a drop-in OpenAI-compatible API, with a privacy-preserving embedding mode that avoids sending prompt content.
Amazon Bedrock Intelligent Prompt Routing
RouterAWS Bedrock's native intelligent prompt router that routes prompts between Anthropic Claude model tiers (Haiku/Sonnet) based on predicted task complexity, with no extra per-routing charge.
Azure AI Foundry Model Router
RouterMicrosoft Azure AI Foundry's native model router that uses a trained ML model to route each prompt in real time to the optimal Azure-hosted model, with Balanced/Cost/Quality mode selection and automatic failover.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(21)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Finance Agent v2 | 53.9 | —Observed 2026-05-28 | — | Source |
| SWE-bench Verified | 88.6 | SWE-bench VerifiedObserved 2026-05-28 | — | Source |
| SWE-bench Pro | 69.2 | SWE-bench ProObserved 2026-05-28 | — | Source |
| Terminal-Bench 2.1 | 74.6 | Terminal-Bench 2.1Observed 2026-05-28 | — | Source |
| Google-Proof Q&A | 93.6 | GPQA DiamondObserved 2026-05-28 | — | Source |
| Humanity's Last Exam — No Tools | 49.8 | HLE no toolsObserved 2026-05-28 | — | Source |
| Humanity's Last Exam — With Tools | 57.9 | HLE with toolsObserved 2026-05-28 | — | Source |
| BrowseComp — Single Agent | 84.3 | BrowseComp single-agentObserved 2026-05-28 | — | Source |
| BrowseComp — Multi-Agent | 88.5 | BrowseComp multi-agent Dynamic WorkflowsObserved 2026-05-28 | — | Source |
| OSWorld-Verified | 83.4 | OSWorld-VerifiedObserved 2026-05-28 | — | Source |
| GDPval-AA | 1890.0 | GDPval-AA ELOObserved 2026-05-28 | — | Source |
| MCP-Atlas | 82.2 | MCP-AtlasObserved 2026-05-28 | — | Source |
| ARC-AGI-2 — High Effort | 72.1 | ARC-AGI-2 high effortObserved 2026-05-28 | — | Source |
| LiveCodeBench | 88.8 | LiveCodeBenchObserved 2026-05-28 | — | Source |
| CursorBench | 63.8 | CursorBench 3.1Observed 2026-06-30 | Configuration: Opus 4.8 Max Harness: CursorBench 3.1 Evaluator: Cursor Confidence: confirmed Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model. | Source |
| GeneBench-Pro | 16.0 | maxObserved 2026-06-30 | — | Source |
| CursorBench | 62.3 | CursorBench 3.2Observed 2026-07-18 | Configuration: Opus 4.8 Max Harness: CursorBench 3.2 productized Cursor-agent workflow Evaluator: Cursor Confidence: confirmed Cost/task: $5.77 Tokens/task: 71,411 Steps/task: 44 Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful. | Source |
| CursorBench | 59.4 | CursorBench 3.2Observed 2026-07-18 | Configuration: Opus 4.8 Extra High Harness: CursorBench 3.2 productized Cursor-agent workflow Evaluator: Cursor Confidence: confirmed Cost/task: $4.50 Tokens/task: 51,121 Steps/task: 40 Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful. | Source |
| CursorBench | 58.0 | CursorBench 3.2Observed 2026-07-18 | Configuration: Opus 4.8 High Harness: CursorBench 3.2 productized Cursor-agent workflow Evaluator: Cursor Confidence: confirmed Cost/task: $3.15 Tokens/task: 33,548 Steps/task: 33 Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful. | Source |
| CursorBench | 56.1 | CursorBench 3.2Observed 2026-07-18 | Configuration: Opus 4.8 Medium Harness: CursorBench 3.2 productized Cursor-agent workflow Evaluator: Cursor Confidence: confirmed Cost/task: $2.81 Tokens/task: 28,384 Steps/task: 32 Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful. | Source |
| CursorBench | 53.1 | CursorBench 3.2Observed 2026-07-18 | Configuration: Opus 4.8 Low Harness: CursorBench 3.2 productized Cursor-agent workflow Evaluator: Cursor Confidence: confirmed Cost/task: $2.02 Tokens/task: 19,624 Steps/task: 27 Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful. | Source |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(4)
Compare Claude Opus 4.8 with other models
Comparison and alternatives
Browse all comparisons →Show all 6 popular comparisonssorted by 7-day search impressions
Cheapest of 6 routes · Anthropic · cache read $0.500