LLM Reference

Best LLMs for Retrieval-Augmented Generation (2026)

Last refreshed 2026-09-03. Next refresh: weekly.

Best LLMs for RAG in 2026 — ranked by context window, retrieval benchmarks, and tool support. Covers document QA to enterprise search.

Verdict

Use GPT-5.6 Sol for RAG today.

GPT-5.6 Terra is the runner-up: 1.05m vs 1.05m on Context.

Researched 57d agoWhy this pickMethodology

How we rank

RAG picks emphasize the strongest sourced long-document benchmark among tracked RAG needles, QA, and retrieval suites, then context window, then recency.

  1. EligibilityModels tagged for the RAG decision task (collections/specialization fit or scores on RULER, ZeroSCROLLs, InfiniteBench, multi-needle, MS MARCO, SQuAD, NaturalQuestions, TriviaQA, etc.).
  2. Primary rankingBest score across the RAG benchmark bundle, then larger declared context window, then newer release.
  3. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. PricingLowest tracked provider input/output where present.
#ModelInput $/1MOutput $/1M
1Nemotron 3 Super-120B-A12B

Signal used: RULER 96.33%

$0.09$0.45
2Llama 4 Scout 17B-16E Instruct
Vision

Signal used: Context 10m

$0.08$0.22
3Gemini 1.5 Pro

Signal used: Context 2m

$1.25$5.00
4GPT-6 Astra
RestrictedReasoningVisionTools

Signal used: Context 1.05m

5GPT-5.6 Sol
ReasoningVisionTools

Signal used: Context 1.05m

$5.00$30.00
6GPT-5.6 Terra
ReasoningVisionTools

Signal used: Context 1.05m

$2.50$15.00
7GPT-5.6 Luna
ReasoningVisionTools

Signal used: Context 1.05m

$1.00$6.00
8GPT-5.5
ReasoningVisionTools

Signal used: Context 1.05m

$5.00$30.00
9GPT-5.5 Pro
ReasoningVisionTools

Signal used: Context 1.05m

$30.00$180.00
10GPT-5.4
ReasoningVisionTools

Signal used: Context 1.05m

$2.50$15.00
11GPT-5.4 Pro
ReasoningVisionTools

Signal used: Context 1.05m

$30.00$180.00
12Muse Spark 1.3
ReasoningVisionTools

Signal used: Context 1.05m

$1.25$4.25
13Muse Spark 1.3 Contributor
ReasoningVisionTools

Signal used: Context 1.05m

$0.10$0.20
14Gemini 3.8 Flash
ReasoningVisionTools

Signal used: Context 1.05m

$0.75$3.75
15Spark X2.5 4B
ReasoningTools

Signal used: Context 1.05m

16Spark X2.5 1.7B
ReasoningTools

Signal used: Context 1.05m

17Hunyuan Hy4 Preview
PreviewReasoningTools

Signal used: Context 1.05m

$0.83$2.50
18Gemini 3.7 Flash
ReasoningVisionTools

Signal used: Context 1.05m

$0.75$3.75
19Muse Spark 1.2
PreviewReasoningVisionTools

Signal used: Context 1.05m

$1.25$4.25
20Gemini 3.6 Flash
ReasoningVisionTools

Signal used: Context 1.05m

$1.50$7.50

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4GPT-5.5

    GPT-5.5 is OpenAI's fully retrained agentic model, released April 23, 2026. Optimised for agentic coding, computer use, knowledge work, and early scientific research. Achieves 82.7% on Terminal-Bench 2.0 (Codex CLI scaffold), 84.9% on GDPval, 58.6% on SWE-Bench Pro, 93.6% on GPQA Diamond, and 82.6% on SWE-Bench Verified (Vals.ai independent harness). Knowledge cutoff December 2025. Supports reasoning effort levels (none/low/medium/high/xhigh). Context window 1,050,000 tokens with a long-context surcharge above 272K tokens. Model ID: gpt-5.5.

    1.05m

    Context

  • #5GPT-5.5 Pro

    GPT-5.5 Pro is OpenAI's premium extra-compute deployment of GPT-5.5, released April 23, 2026. It uses the same underlying weights as GPT-5.5 standard with additional parallel test-time compute for harder tasks. Supports text and image inputs, reasoning effort control, tool use, structured outputs, code execution, a 1,050,000-token context window, and 128K max output. Key datapack rows: Terminal-Bench 2.1 78.2%, SWE-bench Pro 58.6%, GPQA Diamond 93.6%, ARC-AGI-2 high effort 83.3%, BrowseComp Pro compute 90.1%, and FrontierMath Tier 4 39.6%. Official pricing is $30/M input, $180/M output, $10/M batch input, and $45/M batch output; native cached input discount is not listed.

    1.05m

    Context

  • #6GPT-5.4

    GPT-5.4 is OpenAI's flagship frontier reasoning model, released March 5, 2026. It incorporates advances from GPT-5.3-Codex for coding and agentic workflows, and adds 'Thinking' mode with editable reasoning plans. Key capabilities include computer use (navigating interfaces via Playwright), image understanding and generation integration, full-stack web app generation, tool calling, and deep research. Knowledge cutoff is August 31, 2025. Model ID: gpt-5.4.

    1.05m

    Context