LLM Reference

GPT-5.4

Released
2026-03-05
Last refreshed
2026-06-29
Status
Researched 30d ago
ProprietaryCommercial use: conditionalMultimodalCodingRAGAgentsLong contextVisionClassificationJSON / Tool use

GPT-5.4 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 1.05m context window
  • Buyers comparing 4 tracked provider routes

Do not use it for

  • Workloads where another current model has stronger sourced task evidence
Specifications
Family
GPT-5.4
Released
2026-03-05
Context
1.05m
Max output
128,000
Architecture
Decoder Only
Knowledge cutoff
2025-08
Specialization
general
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

Cutting-edge research and development.

San Francisco, California, United States
Founded 2015
Website
Pricing
Output / 1M
$15.00
Input / 1M
$2.50

Cheapest of 4 routes · OpenAI API · cache read $0.250

About

GPT-5.4 is OpenAI's flagship frontier reasoning model, released March 5, 2026. It incorporates advances from GPT-5.3-Codex for coding and agentic workflows, and adds 'Thinking' mode with editable reasoning plans. Key capabilities include computer use (navigating interfaces via Playwright), image understanding and generation integration, full-stack web app generation, tool calling, and deep research. Knowledge cutoff is August 31, 2025. Model ID: gpt-5.4.

GPT-5.4 is a proprietary model. The structured metadata tracks a 1.05m-token context window, multimodal input, reasoning, function calling, tool use, structured outputs, and code execution. This page tracks provider routes through OpenAI API, OpenRouter, Vercel AI Gateway, and 1 more, with the cheapest tracked route listed at $2.5 input and $15 output per 1M tokens. Headline tracked benchmarks include SWE-bench Pro 57.7, Google-Proof Q&A 92.0, and Massive Multi-discipline Multimodal Understanding 82.1.

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ D

2 relevant benchmarks in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

Q/$ D

3 relevant benchmarks in the decision map.

Provider price ladder

Compare all 4

Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MBatch in / outCacheRoute
OpenAI API$2.50$15.00$1.25 / $7.50read $0.250
Serverless
OpenRouter$2.50$15.00-read $0.250
Serverless
Vercel AI Gateway$2.50$15.00-read $0.250
Serverless
AWS Bedrock$2.75$16.50--
Serverless

Available via routers & gateways(16)

Capabilities

VisionMultimodalReasoningFunction CallingTool UseStructured OutputsCode ExecutionPrompt CachingBatch API

Benchmark peer barsfor Coding

Benchmark scores(14)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionSource
SWE-bench Pro57.7SWE-bench Pro Public (pass@1)https://x.com/OpenAIDevs/status/2029620996962242663
Google-Proof Q&A92.0diamondhttps://pricepertoken.com/leaderboards/benchmark/gpqa
Massive Multi-discipline Multimodal Understanding82.1https://mmmu-benchmark.github.io/
MMLU PRO87.5https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro
τ-bench78.3τ-benchhttps://taubench.com/
MultiChallenge69.2MultiChallengehttps://labs.scale.com/leaderboard/multichallenge
SWE-bench Verified71.7SWE-bench Verifiedhttps://artificialanalysis.ai/leaderboards/models
ARC Prize / ARC Challenge73.3ARC-AGI-2https://artificialanalysis.ai/leaderboards/models
Chatbot Arena1479.0Highhttps://arena.ai/leaderboard
MMMU Pro81.2official OpenAI, without tool usehttps://openai.com/index/introducing-gpt-5-4/
ARC-AGI-273.3llm-stats shows 0 (accuracy%)https://llm-stats.com/benchmarks/arc-agi-v2
Humanity's Last Exam41.6HLE (xhigh setting) (accuracy)https://lmcouncil.ai/benchmarks
Terminal-Bench 2.075.1Terminal-Bench 2.0 (accuracy%)https://llm-stats.com/benchmarks/terminal-bench-2
GeneBench-Pro8.9xhighhttps://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf

Migration checks

No linked migration route is available for this model yet.

Compare GPT-5.4 with other models

Show all 69 popular comparisonssorted by 7-day search impressions
GPT-5.4 vs o3 Mini185GPT-5.4 vs GPT-5.2179GPT-5.4 vs Qwen3.6-27B174GPT-5.4 vs Claude Sonnet 4.6173GPT-5.4 vs Claude Sonnet 4.5155GPT-5.4 vs Gemini 2.5 Pro142GPT-5.4 vs Qwen3.6-35B-A3B120GPT-5.4 vs Step 3.5 Flash105GPT-5.4 vs StepFun Step-2100GPT-5.4 vs DeepSeek R190GPT-5.4 vs Together AI Qwen2-72B-Instruct84GPT-5.4 vs Xiaomi MiMo-V2.577GPT-5.4 vs Claude 3.7 Sonnet77GPT-5.4 vs GPT-5.5 Instant59GPT-5.4 vs Ling-2.6-1T52GPT-5.4 vs Qwen3.5-397B-A17B45GPT-5.4 vs Code Davinci 00141GPT-5.4 vs Kimi K2 Thinking37GPT-5.4 vs Qwen2.5-7B-Instruct36GPT-5.4 vs DeepSeek V335GPT-5.4 vs GLM-5 Turbo29GPT-5.4 vs DeepSeek V3.229GPT-5.4 vs Claude Opus 4.629GPT-5.4 vs Tencent Hunyuan Turbo S28GPT-5.4 vs Llama 3.1 70B Instruct27GPT-5.4 vs Llama 3 70B Instruct27GPT-5.4 vs Grok-325GPT-5.4 vs Mistral Large 225GPT-5.4 vs Llama 3.1 405B Instruct25GPT-5.4 vs Claude Opus 4.522GPT-5.4 vs Trinity-Large-Thinking21GPT-5.4 vs Mistral Large 3 675B Instruct21GPT-5.4 vs Llama 2 13B Chat20GPT-5.4 vs Qwen3-235B-A22B18GPT-5.4 vs GLM-5V-Turbo16GPT-5.4 vs Gemini 3 Pro16GPT-5.4 vs Phi-3 Mini 4k15GPT-5.4 vs DeepSeek V3.115GPT-5.4 vs Nano Banana Pro (Gemini 3 Pro Image Preview)14GPT-5.4 vs Nano Banana (Gemini 2.5 Flash Image)14GPT-5.4 vs Llama 3 8B Instruct13GPT-5.4 vs Qwen2.5-72B-Instruct13GPT-5.4 vs GLM-5 9B10GPT-5.4 vs DeepSeek R1 052810GPT-5.4 vs Phi 3.5 Mini Instruct8GPT-5.4 vs Together AI Qwen2-7B-Instruct7GPT-5.4 vs DeepSeek R1 Distill Llama 70B7GPT-5.4 vs Qwen2.5-72B7GPT-5.4 vs Grok 3 Mini7GPT-5.4 vs Kimi K2.56GPT-5.4 vs Qwen3.5-35B-A3B6GPT-5.4 vs Gemini 2.5 Flash Live API6GPT-5.4 vs Kimi K2 Instruct6GPT-5.4 vs Qwen3.5-122B-A10B5GPT-5.4 vs Phi-4 Reasoning Vision 15B5GPT-5.4 vs Gemini 2.5 Pro Preview 05-065GPT-5.4 vs Mistral Nemotron5GPT-5.4 vs Qwen3-9B4GPT-5.4 vs Gemma 7B Instruct4GPT-5.4 vs Magistral Small 25064GPT-5.4 vs Mixtral 8x7B3GPT-5.4 vs Mistral Magistral Small 25093GPT-5.4 vs Llama 3.2 1B Instruct2GPT-5.4 vs Qwen3-Max1GPT-5.4 vs ShieldGemma 9B1GPT-5.4 vs Qwen2-7B-Instruct1GPT-5.4 vs Qwen3.5-27B1GPT-5.4 vs Kimi K2 Thinking Turbo1GPT-5.4 vs Phi-4 Mini Flash Reasoning1

Frequently asked questions

What is the context window of GPT-5.4?

GPT-5.4 has a context window of 1.05m tokens.

What is the max output of GPT-5.4?

GPT-5.4 can generate up to 128,000 output tokens.

How much does GPT-5.4 cost?

GPT-5.4 pricing ranges from $2.50/1M to $2.75/1M input tokens depending on the provider.

When was GPT-5.4 released?

GPT-5.4 was released on 2026-03-05.

Which providers offer GPT-5.4?

GPT-5.4 is available from 4 providers: OpenAI API, OpenRouter, Vercel AI Gateway, AWS Bedrock.

What benchmarks has GPT-5.4 been tested on?

GPT-5.4 has been evaluated on 14 benchmarks, including SWE-bench Pro, Google-Proof Q&A, Massive Multi-discipline Multimodal Understanding, MMLU PRO, τ-bench.