GPT-5.5
GPT-5.5 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 1.05m context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- GPT-5.5
- Released
- 2026-04-23
- Context
- 1.05m
- Max output
- 128,000
- Architecture
- Decoder Only
- Knowledge cutoff
- 2025-12
- Specialization
- general
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
Cheapest of 4 routes · OpenAI API · cache read $0.500
About
GPT-5.5 is OpenAI's fully retrained agentic model, released April 23, 2026. Optimised for agentic coding, computer use, knowledge work, and early scientific research. Achieves 82.7% on Terminal-Bench 2.0 (Codex CLI scaffold), 84.9% on GDPval, 58.6% on SWE-Bench Pro, 93.6% on GPQA Diamond, and 82.6% on SWE-Bench Verified (Vals.ai independent harness). Knowledge cutoff December 2025. Supports reasoning effort levels (none/low/medium/high/xhigh). Context window 1,050,000 tokens with a long-context surcharge above 272K tokens. Model ID: gpt-5.5.
GPT-5.5 is a proprietary model. The structured metadata tracks a 1.05m-token context window, multimodal input, reasoning, function calling, tool use, structured outputs, and code execution. This page tracks provider routes through OpenAI API, OpenRouter, Vercel AI Gateway, and 1 more, with the cheapest tracked route listed at $5 input and $30 output per 1M tokens. Headline tracked benchmarks include Google-Proof Q&A 93.6, SWE-bench Pro 58.6, and Chatbot Arena 1488.0.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ D4 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ D1 relevant benchmark in the decision map.
Provider price ladder
Compare all 4Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Batch in / out | Cache | Route |
|---|---|---|---|---|---|
| OpenAI API | $5.00 | $30.00 | $2.50 / $15.00 | read $0.500 | Serverless |
| OpenRouter | $5.00 | $30.00 | - | - | Serverless |
| Vercel AI Gateway | $5.00 | $30.00 | - | read $0.500 | Serverless |
| AWS Bedrock | $5.50 | $33.00 | - | - | Serverless |
Available via routers & gateways(16)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
AIRouter
RouterCommercial LLM router that analyzes incoming requests and routes to the optimal model for cost/quality/latency via a drop-in OpenAI-compatible API, with a privacy-preserving embedding mode that avoids sending prompt content.
Amazon Bedrock Intelligent Prompt Routing
RouterAWS Bedrock's native intelligent prompt router that routes prompts between Anthropic Claude model tiers (Haiku/Sonnet) based on predicted task complexity, with no extra per-routing charge.
Helicone
GatewayObservability-first AI gateway with routing, caching, rate limiting, and request tracing; Apache 2.0 open-source core with a managed hosted tier for logging and analytics.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(26)
| Benchmark | Score | Version | Source |
|---|---|---|---|
| Google-Proof Q&A | 93.6 | diamond | https://openai.com/index/introducing-gpt-5-5/ |
| SWE-bench Pro | 58.6 | SWE-bench Pro (pass@1) | https://www.vellum.ai/blog/everything-you-need-to-know-about-gpt-5-5 |
| Chatbot Arena | 1488.0 | High | https://arena.ai/leaderboard |
| MMLU PRO | 88.1 | — | https://openai.com/index/introducing-gpt-5-5/ |
| MMMU Pro | 88.3 | Vals.ai standardized CoT harness | https://www.vals.ai/benchmarks/mmmu |
| HumanEval | 94.2 | — | https://openai.com/index/introducing-gpt-5-5/ |
| Instruction-Following Evaluation | 92.1 | — | https://openai.com/index/introducing-gpt-5-5/ |
| SWE-bench Verified | 82.6 | SWE-bench Verified | https://www.vals.ai/benchmarks/swebench |
| Massive Multitask Language Understanding | 92.4 | MMLU (accuracy) | https://tokenmix.ai/blog/gpt-5-5-spud-review-88-swe-bench-2026 |
| Terminal-Bench 2.0 | 82.7 | Terminal-Bench 2.0 (accuracy%) | https://llm-stats.com/benchmarks/terminal-bench-2 |
| Aider Polyglot | 88.0 | Listed as 'gpt-5 (high)' = 88 (percent_correct) | https://aider.chat/docs/leaderboards/ |
| AIME 2025 | 81.2 | AIME 2025 (accuracy) | https://codersera.com/blog/openai-may-2026-updates-roundup/ |
| ARC-AGI-2 | 85.0 | ARC-AGI-2 (accuracy%) | https://benchlm.ai/benchmarks/arcAgi2 |
| BrowseComp | 84.4 | BrowseComp (accuracy%) | https://benchlm.ai/benchmarks/browseComp |
| Humanity's Last Exam | 41.4 | HLE without tools (accuracy) | https://www.vellum.ai/blog/everything-you-need-to-know-about-gpt-5-5 |
| MCP-Atlas | 75.3 | MCP-Atlas (accuracy%) | https://benchlm.ai/benchmarks/mcpAtlas |
| GDPval | 84.9 | GDPval official launch score | https://openai.com/index/introducing-gpt-5-5/ |
| Terminal-Bench 2.1 | 78.2 | Terminal-Bench 2.1 | https://www.morphllm.com/best-ai-model-for-coding |
| OSWorld-Verified | 78.7 | OSWorld-Verified | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| GDPval-AA | 1769.0 | GDPval-AA ELO | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| Blueprint-Bench 2 | 36.2 | Blueprint-Bench 2 | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| AutomationBench | 12.9 | AutomationBench | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| Legal Agent Benchmark | 2.1 | Legal Agent Benchmark | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| GDP.pdf | 24.9 | GDP.pdf vision, no tools | https://www.anthropic.com/news/claude-fable-5-mythos-5 |
| CursorBench | 64.3 | CursorBench 3.1 | https://cursor.com/evals |
| GeneBench-Pro | 12.0 | xhigh | https://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf |
Migration checks
Rankings & picks(10)
Compare GPT-5.5 with other models
Comparison and alternatives
Browse all comparisons →Show all 78 popular comparisonssorted by 7-day search impressions
Frequently asked questions
What is the context window of GPT-5.5?
GPT-5.5 has a context window of 1.05m tokens.
What is the max output of GPT-5.5?
GPT-5.5 can generate up to 128,000 output tokens.
How much does GPT-5.5 cost?
GPT-5.5 pricing ranges from $5/1M to $5.50/1M input tokens depending on the provider.
When was GPT-5.5 released?
GPT-5.5 was released on 2026-04-23.
Which providers offer GPT-5.5?
GPT-5.5 is available from 4 providers: OpenAI API, OpenRouter, Vercel AI Gateway, AWS Bedrock.
What benchmarks has GPT-5.5 been tested on?
GPT-5.5 has been evaluated on 26 benchmarks, including Google-Proof Q&A, SWE-bench Pro, Chatbot Arena, MMLU PRO, MMMU Pro.
Cheapest of 4 routes · OpenAI API · cache read $0.500