The 18 LLM leaderboards we'd actually use.
Editor-curated picks for every job — coding agents, CLI workflows, writing, image, voice — benchmark- and pricing-aware, refreshed weekly.
Developers
Coding
Anthropic's new flagship: 80.3% SWE-bench Pro, 96% SWE-bench Verified on Vals.ai, and 85.0% OSWorld-Verified make it the best production coding pick for non-trivial engineering tasks.
Agents
Everyday agent at Sonnet list $3 / $15 per 1M tokens, with a 1M window, 81.2% OSWorld-Verified, and 86.6% BrowseComp multi-agent.
Tool use
Current Sonnet for function calling and structured outputs: 1M window, 86.6% BrowseComp multi-agent, 81.2% OSWorld-Verified, list $3 / $15 per 1M tokens. No BFCL row on the live model page.
Open weights
Best open-weights model we've tested: #1 LiveCodeBench (93.5), 80.6 SWE-bench, 1M context, $0.87 out.
Long context
1M-token window at Sonnet list $3 / $15 per 1M tokens, with 86.6% BrowseComp multi-agent and 81.2% OSWorld-Verified when the job is tools across a full corpus.
Cheap
$0.12 / 1M out (live table $0.1173) with LiveCodeBench 91.6 and a 1M window, open source.
Knowledge workers
Writing
Tops Chatbot Arena (1503) and writes paragraphs you'd ship; understands tone notes and edits like a copy chief.
Research
GDPval-AA ELO 1932 and Anthropic-reported finance, trading, and analytics wins make it the strongest general knowledge-work pick; do not use Mythos-only HLE rows as Fable evidence.
Summarization
1M context plus MMLU-Pro 88.6 at $3 out — handles a 500-page transcript faithfully without breaking the budget.
Docs Q&A
Highest groundedness and graceful refusal — says "not in the docs" instead of guessing.
Translation
Best coverage of lower-resource languages with strong idiom handling and a huge context for document-level consistency.
Data & SQL
Reads large schemas (1.05M context) and writes SQL that picks the right join and respects dialect quirks, at $5 / $30 per 1M tokens.
Creatives
Image
FLUX.2 Klein 9B is a 9B distilled text-to-image model on fal. Not FLUX.2 Dev.
Video
Best overall video quality in the catalog: 30-second clips, native audio, and up to 4K through Vertex AI.
Voice (TTS)
ElevenLabs' GA TTS (February 2, 2026), API id `eleven_v3`: expressive multi-speaker speech in 70+ languages at $0.10 per 1K characters.
Transcription
Lowest WER on noisy real-world audio with the broadest language coverage; cheap to self-host.
Music
Most production-ready output; stems are usable and the vocals are believable.
Image editing
Best inpaint preservation — brand colors and untouched regions stay put across edits.
- 1 · EligibilityA model is eligible for a board only if it's tagged with that use case. Editors pin a handful per board, with exactly one designated Editor's Choice.
- 2 · Editorial tiersEach pick is bucketed into one of three qualitative tiers — Excellent · Strong · Solid. No decimals, no composite score, just editorial judgment.
- 3 · Cross-checkPicks are opinionated. /best → is the objective benchmark composite for the same capability.