LLM Reference

GPT-5.6 Sol

Released
2026-07-09
Last refreshed
2026-07-09
Status
Researched 39d ago
ProprietaryCommercial use: conditionalMultimodalCodingRAGAgentsLong contextVisionJSON / Tool use

GPT-5.6 Sol is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 1.05m context window
  • Buyers comparing 2 tracked provider routes

Do not use it for

  • Workloads where another current model has stronger sourced task evidence
Specifications
Family
GPT-5.6
Released
2026-07-09
Context
1.05m
Max output
128,000
Architecture
Decoder Only
Specialization
general
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

Cutting-edge research and development.

San Francisco, California, United States
Founded 2015
Website
Pricing
Output / 1M
$30.00
Input / 1M
$5.00

Cheapest of 2 routes · OpenAI API · cache read $0.500

About

OpenAI's flagship GPT-5.6 model and highest-capability tier in the Sol, Terra, and Luna naming system. GPT-5.6 Sol is built for demanding reasoning, long-horizon coding, agentic workflows, and cybersecurity tasks, introducing max reasoning effort and ultra multi-agent mode. Generally available July 9, 2026 across ChatGPT, Codex, and the OpenAI API (model IDs gpt-5.6-sol and alias gpt-5.6). Supports text and image inputs, reasoning, tool use, code execution, prompt caching, and Batch API with a 1,050,000-token context window and 128K max output.

GPT-5.6 Sol is a proprietary model in the GPT-5.6 family. The structured metadata tracks a 1.05m-token context window, multimodal input, reasoning, function calling, tool use, and code execution. This page tracks provider routes through OpenAI API and OpenRouter, with the cheapest tracked route listed at $5 input and $30 output per 1M tokens. Headline tracked benchmarks include HealthBench Professional 60.5, HealthBench 57.0, and HealthBench Hard 33.1.

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ D

1 relevant benchmark in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare all 2

Compare API pricing across 2 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MBatch in / outCacheRoute
OpenAI API$5.00$30.00$2.50 / $15.00read $0.500
Serverless
OpenRouter$5.00$30.00--
Serverless

Available via routers & gateways(15)

Capabilities

VisionMultimodalReasoningFunction CallingTool UseCode ExecutionPrompt CachingBatch API

Benchmark peer barsfor Coding

Benchmark scores(21)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
HealthBench Professional60.5HealthBench Professional; official OpenAI scoring approachObserved 2026-07-09Source
HealthBench57.0length-adjusted; raw 55.6; mean response 1764 charsObserved 2026-06-26Source
HealthBench Hard33.1length-adjusted; raw 31.1; mean response 1751 charsObserved 2026-06-26Source
HealthBench Consensus95.5length-adjusted; raw 95.3; mean response 1740 charsObserved 2026-06-26Source
Terminal-Bench 2.188.8Terminal-Bench 2.1; standard mode; OpenAI GA launch harnessObserved 2026-07-09Source
GeneBench-Pro31.5Pro mode enabled / GPT Pro (Extended); max reasoning; GA launch cites 28.7% standard modeObserved 2026-07-09Source
DeepSWE 1.172.7DeepSWE 1.1; OpenAI GA launch tableObserved 2026-07-09Source
BrowseComp90.4BrowseComp; standard Sol; agentic browsing; OpenAI GA launch harnessObserved 2026-07-09Source
OSWorld62.6OSWorld 2.0; computer useObserved 2026-07-09Source
Google-Proof Q&A94.6GPQA Diamond; max reasoningObserved 2026-07-09Source
MMMU Pro83.0MMMU Pro; no tools; max reasoningObserved 2026-07-09Source
Agents' Last Exam52.7OpenAI launch table; professional long-horizon workflows; standard modeObserved 2026-07-09Source
Artificial Analysis Intelligence Index58.9Artificial Analysis Intelligence Index v4.1; max reasoningObserved 2026-07-09Source
Artificial Analysis Coding Agent Index80.0Artificial Analysis Coding Agent Index v1.1; max reasoningObserved 2026-07-09Source
SWE-bench Pro64.6SWE-Bench Pro; OpenAI GA launch table; codingObserved 2026-07-09Source
BrowseComp — Multi-Agent92.2BrowseComp; Sol Ultra multi-agent mode; agentic browsing; non-comparable to standard single-model rowsObserved 2026-07-09Source
CursorBench67.2CursorBench 3.2Observed 2026-07-18
Configuration: GPT-5.6 Sol Max
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $5.69
Tokens/task: 28,320
Steps/task: 48
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source
CursorBench64.5CursorBench 3.2Observed 2026-07-18
Configuration: GPT-5.6 Sol Extra High
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $3.88
Tokens/task: 19,699
Steps/task: 38
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source
CursorBench63.5CursorBench 3.2Observed 2026-07-18
Configuration: GPT-5.6 Sol High
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $2.79
Tokens/task: 13,867
Steps/task: 32
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source
CursorBench60.0CursorBench 3.2Observed 2026-07-18
Configuration: GPT-5.6 Sol Medium
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $1.95
Tokens/task: 9,747
Steps/task: 27
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source
CursorBench52.6CursorBench 3.2Observed 2026-07-18
Configuration: GPT-5.6 Sol Low
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $1.01
Tokens/task: 5,104
Steps/task: 19
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source

Migration checks

No linked migration route is available for this model yet.

API versions

gpt-5.6

Frequently asked questions

What is the context window of GPT-5.6 Sol?

GPT-5.6 Sol has a context window of 1.05m tokens.

What is the max output of GPT-5.6 Sol?

GPT-5.6 Sol can generate up to 128,000 output tokens.

How much does GPT-5.6 Sol cost?

GPT-5.6 Sol is available at $5.00/1M input tokens through OpenAI API.

When was GPT-5.6 Sol released?

GPT-5.6 Sol was released on 2026-07-09.

Which providers offer GPT-5.6 Sol?

GPT-5.6 Sol is available from 2 providers: OpenAI API, OpenRouter.

What benchmarks has GPT-5.6 Sol been tested on?

GPT-5.6 Sol has been evaluated on 21 benchmarks, including HealthBench Professional, HealthBench, HealthBench Hard, HealthBench Consensus, Terminal-Bench 2.1.