LLM Reference

The long context leaderboard · for developers

Best for long context

Decision area
Past 200K tokens without falling apart.
Editor picks
6 editor picks
Eligible models
13 eligible models
See raw /best
EDITOR'S CHOICEResearched 8d ago

Claude Sonnet 5

Anthropic · 1m context
Excellent

The models still credible when the window is genuinely full.

1M-token window at Sonnet list $3 / $15 per 1M tokens, with 86.6% BrowseComp multi-agent and 81.2% OSWorld-Verified when the job is tools across a full corpus.

The numbers
$/1M out
$10.00
$2.00 input
Context
1m
max window
Pros
  • +86.6% BrowseComp multi-agent
  • +81.2% OSWorld-Verified
  • +1M context at Sonnet list price
Cons
  • $15 / 1M out at list

Also worth picking

The runners-up

ranked by editorial pick orderEditorial tiersExcellentStrongSolid
Google DeepMind · 1m
$5.00 / 1M out
The long-context standard — 1M tokens with cheap input that keeps the bill survivable.
Anthropic · 1m
$25.00 / 1M out
Editorial candidate for complex agentic work across a 1M window with 128K synchronous output. Anthropic includes the full window at standard pricing; no retrieval-specific ranking claim is implied.
Anthropic · 1m
$25.00 / 1M out
Catches subtle cross-references across a full 1M window better than anyone; pricey but precise.
OpenAI · 1.05m
$15.00 / 1M out
Same 1.05M window as GPT-5.5 at half the output price ($15 vs $30) — the value choice for long analytical passes.
DeepSeek · 1m
$0.12 / 1M out
1M context at $0.12 out, open weights — the value pick for huge inputs.

Eligibility

13 models are eligible for this board

Eligibility means tagged with useCases: [long-context]. Pins must come from this pool.
All picks