The agents leaderboard · for developers

Best for agents

Decision area
Tool-use loops that don't crash on turn 12.
Editor picks
7 editor picks
Eligible models
67 eligible models
See raw /best
EDITOR'S CHOICEResearched 31d ago

Claude Sonnet 5

Anthropic · 1m context
Excellent

The most agentic Sonnet yet — strong computer use and multi-step recovery at Sonnet pricing.

Everyday agent at Sonnet list $3 / $15 per 1M tokens, with a 1M window, 81.2% OSWorld-Verified, and 86.6% BrowseComp multi-agent.

The numbers
$/1M out
$10.00
$2.00 input
Context
1m
max window
Pros
  • +81.2% OSWorld-Verified
  • +86.6% BrowseComp multi-agent
  • +1M context at Sonnet list price
Cons
  • −No τ-bench or MultiChallenge row yet
  • −$15 / 1M out at list

Also worth picking

The runners-up

ranked by editorial pick orderEditorial tiersExcellentStrongSolid
Anthropic · 1m
$25.00 / 1M out
Editorial candidate for long-horizon agents: 1M context, native tool use and structured outputs, plus vendor-reported Frontier-Bench and coding-agent rows. Do not claim τ-bench or BFCL leadership without compatible evidence.
Anthropic · 1m
$15.00 / 1M out
Best generally-available τ-bench (87.5) among predecessors — solid fallback until Sonnet 5 publishes comparable agent-bench rows.
Anthropic · 1m
$50.00 / 1M out
OSWorld-Verified 85.0% and 80.3% SWE-bench Pro make it a top-tier computer-use and coding-agent candidate, but no Fable-specific τ-bench score is published yet.
Zhipu AI · 200k
$2.08 / 1M out
τ-bench 82.1, open weights at $2.08 out — the best price-per-step agent we'd run always-on.
OpenAI · 1.05m
$15.00 / 1M out
Picked over GPT-5.5 here: it carries the measured τ-bench score (78.3) and runs at half the output price ($15 vs $30) — decisive for always-on agents.
Google DeepMind · 1m
$5.00 / 1M out
Best for chatty agents with many read-only tools and large contexts.

Eligibility

67 models are eligible for this board

Eligibility means tagged with useCases: [agents]. Pins must come from this pool.
All picks