GPT-5.4
GPT-5.4 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 1.05m context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- GPT-5.4
- Released
- 2026-03-05
- Context
- 1.05m
- Max output
- 128,000
- Architecture
- Decoder Only
- Knowledge cutoff
- 2025-08
- Specialization
- general
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
Cheapest of 4 routes · OpenAI API · cache read $0.250
About
GPT-5.4 is OpenAI's flagship frontier reasoning model, released March 5, 2026. It incorporates advances from GPT-5.3-Codex for coding and agentic workflows, and adds 'Thinking' mode with editable reasoning plans. Key capabilities include computer use (navigating interfaces via Playwright), image understanding and generation integration, full-stack web app generation, tool calling, and deep research. Knowledge cutoff is August 31, 2025. Model ID: gpt-5.4.
GPT-5.4 is a proprietary model. The structured metadata tracks a 1.05m-token context window, multimodal input, reasoning, function calling, tool use, structured outputs, and code execution. This page tracks provider routes through OpenAI API, OpenRouter, Vercel AI Gateway, and 1 more, with the cheapest tracked route listed at $2.5 input and $15 output per 1M tokens. Headline tracked benchmarks include SWE-bench Pro 57.7, Google-Proof Q&A 92.0, and Massive Multi-discipline Multimodal Understanding 82.1.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ D2 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ D3 relevant benchmarks in the decision map.
Provider price ladder
Compare all 4Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Batch in / out | Cache | Route |
|---|---|---|---|---|---|
| OpenAI API | $2.50 | $15.00 | $1.25 / $7.50 | read $0.250 | Serverless |
| OpenRouter | $2.50 | $15.00 | - | read $0.250 | Serverless |
| Vercel AI Gateway | $2.50 | $15.00 | - | read $0.250 | Serverless |
| AWS Bedrock | $2.75 | $16.50 | - | - | Serverless |
Available via routers & gateways(16)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
AIRouter
RouterCommercial LLM router that analyzes incoming requests and routes to the optimal model for cost/quality/latency via a drop-in OpenAI-compatible API, with a privacy-preserving embedding mode that avoids sending prompt content.
Amazon Bedrock Intelligent Prompt Routing
RouterAWS Bedrock's native intelligent prompt router that routes prompts between Anthropic Claude model tiers (Haiku/Sonnet) based on predicted task complexity, with no extra per-routing charge.
Helicone
GatewayObservability-first AI gateway with routing, caching, rate limiting, and request tracing; Apache 2.0 open-source core with a managed hosted tier for logging and analytics.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(14)
| Benchmark | Score | Version | Source |
|---|---|---|---|
| SWE-bench Pro | 57.7 | SWE-bench Pro Public (pass@1) | https://x.com/OpenAIDevs/status/2029620996962242663 |
| Google-Proof Q&A | 92.0 | diamond | https://pricepertoken.com/leaderboards/benchmark/gpqa |
| Massive Multi-discipline Multimodal Understanding | 82.1 | — | https://mmmu-benchmark.github.io/ |
| MMLU PRO | 87.5 | — | https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro |
| τ-bench | 78.3 | τ-bench | https://taubench.com/ |
| MultiChallenge | 69.2 | MultiChallenge | https://labs.scale.com/leaderboard/multichallenge |
| SWE-bench Verified | 71.7 | SWE-bench Verified | https://artificialanalysis.ai/leaderboards/models |
| ARC Prize / ARC Challenge | 73.3 | ARC-AGI-2 | https://artificialanalysis.ai/leaderboards/models |
| Chatbot Arena | 1479.0 | High | https://arena.ai/leaderboard |
| MMMU Pro | 81.2 | official OpenAI, without tool use | https://openai.com/index/introducing-gpt-5-4/ |
| ARC-AGI-2 | 73.3 | llm-stats shows 0 (accuracy%) | https://llm-stats.com/benchmarks/arc-agi-v2 |
| Humanity's Last Exam | 41.6 | HLE (xhigh setting) (accuracy) | https://lmcouncil.ai/benchmarks |
| Terminal-Bench 2.0 | 75.1 | Terminal-Bench 2.0 (accuracy%) | https://llm-stats.com/benchmarks/terminal-bench-2 |
| GeneBench-Pro | 8.9 | xhigh | https://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(10)
Compare GPT-5.4 with other models
Comparison and alternatives
Browse all comparisons →Show all 69 popular comparisonssorted by 7-day search impressions
Frequently asked questions
What is the context window of GPT-5.4?
GPT-5.4 has a context window of 1.05m tokens.
What is the max output of GPT-5.4?
GPT-5.4 can generate up to 128,000 output tokens.
How much does GPT-5.4 cost?
GPT-5.4 pricing ranges from $2.50/1M to $2.75/1M input tokens depending on the provider.
When was GPT-5.4 released?
GPT-5.4 was released on 2026-03-05.
Which providers offer GPT-5.4?
GPT-5.4 is available from 4 providers: OpenAI API, OpenRouter, Vercel AI Gateway, AWS Bedrock.
What benchmarks has GPT-5.4 been tested on?
GPT-5.4 has been evaluated on 14 benchmarks, including SWE-bench Pro, Google-Proof Q&A, Massive Multi-discipline Multimodal Understanding, MMLU PRO, τ-bench.
Cheapest of 4 routes · OpenAI API · cache read $0.250