LLM Reference

Best LLMs for Writing (2026)

Last refreshed 2026-09-02. Next refresh: weekly.

The best LLMs for writing in 2026, ranked by human preference. Covers long-form essays, creative prose, and marketing copy — with pricing.

Verdict

Use Muse Spark 1.3 for long-form writing today.

Muse Spark 1.3 Contributor is the runner-up: — vs — on Arena.

Researched todayWhy this pickMethodology

How we rank

Writing picks are for essays, drafts, and long-form prose. The target rubric is EQ-Bench Creative Writing v3 when coverage is available; the live fallback is Chatbot Arena, then MMLU.

  1. EligibilityGeneral chat models excluding pure code/embedding SKUs.
  2. Target rankingEQ-Bench Creative Writing v3 becomes the primary signal once it exists in seed data for enough current models.
  3. Current live rankingUntil that coverage lands, the page keeps the current Chatbot Arena preference score fallback, then MMLU, then newer release.
  4. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. Why this ranking?Arena is an imperfect but useful human-preference proxy for prose quality; MMLU is only a general capability floor, not a style or brand-voice metric.
  6. Marketing boundaryUse `/best/marketing` for ad copy, email, social posts, localization, and brand-voice content; this page stays focused on general writing.
#ModelInput $/1MOutput $/1M
1Claude Opus 4.7
ReasoningVisionTools

Arena: 1503

$5.00$25.00
2Claude Opus 4.6
ReasoningVisionTools

Arena: 1501

$5.00$25.00
3Gemini 3.1 Pro Preview
PreviewVisionTools

Arena: 1493

$2.00$12.00
4Muse Spark
ReasoningVisionTools

Arena: 1491

5GPT-5.5
ReasoningVisionTools

Arena: 1488

$5.00$30.00
6Gemini 3 Pro
VisionTools

Arena: 1486

$1.25$5.00
7GPT-5.4
ReasoningVisionTools

Arena: 1479

$2.50$15.00
8ERNIE 5.1
Tools

Arena: 1476

$0.59$2.65
9Qwen3.7-Max
ReasoningTools

Arena: 1475

$1.25$3.75
10GLM-5.1
ReasoningTools

Arena: 1475

$1.05$3.50
11Gemini 3 Flash
PreviewVisionTools

Arena: 1467

$0.50$3.00
12Claude Opus 4.5
ReasoningVisionTools

Arena: 1466

$5.00$25.00
13Grok 4.1
ReasoningVisionTools

Arena: 1464

14Claude Sonnet 4.6
ReasoningVisionTools

Arena: 1459

$3.00$15.00
15DeepSeek V4 Pro
ReasoningTools

Arena: 1456

$0.43$0.87
16DeepSeek V4 Flash
ReasoningTools

Arena: 1437

$0.06$0.12
17Gemini 3.1 Flash-Lite
VisionTools

Arena: 1432

$0.25$1.50
18o3
ReasoningVisionTools

Arena: 1412

$2.00$8.00
19Gemini 2.5 Pro
ReasoningVisionTools

Arena: 1398

$1.25$10.00
20DeepSeek R1
Reasoning

Arena: 1372

$0.10$0.30

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Qwen3.8-Max-0902 is Alibaba's dated API snapshot of Qwen3.8-Max, listed September 2, 2026 on QwenCloud and Model Studio as model ID qwen3.8-max-0902 (first-party alias qwen3.8-max-2026-09-02). Official copy: an upgraded snapshot of qwen3.8-max with stronger coding (engineering-scale / long-horizon autonomous development), collaborative-agent multi-tool orchestration, and refined native vision (chart reasoning, document parsing, multimodal perception). Retains the 1,000,000-token context window, thinking mode, and full tool ecosystem. Input image/text/video, output text; function calling, structured outputs, context cache. International QwenCloud / Model Studio Singapore list price $2 input / $6 output per 1M tokens; implicit cache $0.25; explicit cache create $2.50; explicit cache read $0.17. No first-party 0902 weight drop. Sibling of seeded qwen3.8-max under family qwen3.8. Compare it for Agents and Coding.

    Arena

  • Anthropic's generally available Mythos-class model for demanding reasoning, long-horizon agents, and ambitious coding, released September 1, 2026. Same underlying model as Claude Mythos 5.1 with cybersecurity and biology safeguards (flagged queries route to Opus 4.8 for cyber and Opus 5 for biology; those reroutes are not billed at Fable prices). Claude API id claude-fable-5-1. List price $10 per 1M input tokens and $50 per 1M output tokens; cache reads $0.25 per 1M (0.025x base input, 75% less than Fable 5's $1.00). 1M-token context, 128K max output, adaptive thinking always on (default effort high), text and image input to text output, vision, tool use. Available to Pro, Max, Team, and Enterprise users and on the Claude Platform natively, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Reliable knowledge cutoff June 2026. Sibling of seeded claude-fable-5 under family claude-fable. Compare it for Agents and Coding.

    Arena

  • #6Claude Mythos 5.1Restricted

    Anthropic's invitation-only Mythos-class model, released September 1, 2026. Same underlying model as Claude Fable 5.1 without the cybersecurity and biology safeguards. Access is limited to a small set of vetted organizations through trusted access / Project Glasswing; Anthropic currently makes it available to a set of US organizations and is working to expand. Claude API id claude-mythos-5-1. Pricing starts at $10 per 1M input tokens and $50 per 1M output tokens (limited availability; no cheaper public tier published). Cache reads $0.25 per 1M. 1M-token context, 128K max output, adaptive thinking always on, text and image input to text output. Claude Security now runs on Mythos 5.1. Sibling of seeded claude-mythos-5 under family claude-mythos.

    Arena