LLM Reference

Best LLMs for Retrieval-Augmented Generation (2026)

Last refreshed 2026-09-21. Next refresh: weekly.

Best LLMs for RAG in 2026 — ranked by context window, retrieval benchmarks, and tool support. Covers document QA to enterprise search.

Verdict

Use Arrow 2 Telos for RAG today.

GPT-6 Astra is the runner-up: 1.05m vs 1.05m on Context.

Researched 4d agoWhy this pickMethodology

How we rank

RAG picks emphasize the strongest sourced long-document benchmark among tracked RAG needles, QA, and retrieval suites, then context window, then recency.

  1. Eligibility — Models tagged for the RAG decision task (collections/specialization fit or scores on RULER, ZeroSCROLLs, InfiniteBench, multi-needle, MS MARCO, SQuAD, NaturalQuestions, TriviaQA, etc.).
  2. Primary ranking — Best score across the RAG benchmark bundle, then larger declared context window, then newer release.
  3. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. Pricing — Lowest tracked provider input/output where present.
#ModelInput $/1MOutput $/1M
1Nemotron 3 Super-120B-A12B

Signal used: RULER 96.33%

$0.09$0.45
2Llama 4 Scout 17B-16E Instruct
Vision

Signal used: Context 10m

$0.08$0.22
3Gemini 1.5 Pro

Signal used: Context 2m

$1.25$5.00
4Arrow 2 Telos
VisionTools

Signal used: Context 1.05m

$6.00$30.00
5GPT-6 Astra
RestrictedReasoningVisionTools

Signal used: Context 1.05m

$10.00$50.00
6GPT-5.6 Luna
ReasoningVisionTools

Signal used: Context 1.05m

$1.00$6.00
7GPT-5.6 Sol
ReasoningVisionTools

Signal used: Context 1.05m

$5.00$30.00
8GPT-5.6 Terra
ReasoningVisionTools

Signal used: Context 1.05m

$2.50$15.00
9GPT-5.5
ReasoningVisionTools

Signal used: Context 1.05m

$5.00$30.00
10GPT-5.5 Pro
ReasoningVisionTools

Signal used: Context 1.05m

$30.00$180.00
11GPT-5.4
ReasoningVisionTools

Signal used: Context 1.05m

$2.50$15.00
12GPT-5.4 Pro
ReasoningVisionTools

Signal used: Context 1.05m

$30.00$180.00
13Xiaomi MiMo-V2.6-Pro
ReasoningVisionTools

Signal used: Context 1.05m

$0.43$0.87
14Nex-N2.5-Max
ReasoningTools

Signal used: Context 1.05m

——
15Gemini 3.8 Flash
ReasoningVisionTools

Signal used: Context 1.05m

$0.75$3.75
16Muse Spark 1.3
ReasoningVisionTools

Signal used: Context 1.05m

$1.25$4.25
17Muse Spark 1.3 Contributor
ReasoningVisionTools

Signal used: Context 1.05m

$0.10$0.20
18Spark X2.5 1.7B
ReasoningTools

Signal used: Context 1.05m

——
19Spark X2.5 4B
ReasoningTools

Signal used: Context 1.05m

——
20Hunyuan Hy4 Preview
PreviewReasoningTools

Signal used: Context 1.05m

$0.83$2.50

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Nex-N2.5-Max is Nex AGI's trillion-scale text-only MoE agentic model in the Nex-N2.5 family, released September 8, 2026. Apache-2 weights on Hugging Face (nex-agi/Nex-N2.5-Max). 1.6T total / 49B active MoE with a native 1,048,576-token context window (config max_position_embeddings). Supports reasoning, function calling, and tool use; no vision or multimodal inputs. No OpenRouter listing in this seed pass. Sibling Nex-N2.5-mini is the multimodal MoE variant; Nex-N2.5-Pro is held from this seed pass.

    1.05m

    Context

  • Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8. Compare it for Coding, Agents, Long context, Vision, and JSON / Tool use.

    1.05m

    Context

  • Muse Spark 1.3 is Meta Superintelligence Labs' latest Muse Spark checkpoint, released September 2, 2026. Official Meta Model API model ID: muse-spark-1.3. First-party models docs: the latest version, recommended for new work, tuned for agentic workflows (multi-step tool, browser, and long-horizon tasks) with improved coding over 1.2; default model in docs examples. Product line (dev.meta.ai): Built for long-horizon coding and agentic work. Cleaner code, fewer tokens. First-party research post: rolling out today in Muse Code and Meta Model API; previously available reasoning modes available today with max reasoning coming shortly after additional safety testing. Official table: text, image, video, audio*, PDF input and text output; 1,048,576-token context. Official quickstart limits: context 1048576, output 131072; input modalities text/image/pdf/video, output text. Audio understanding in 1.3 is currently not fully supported (use 1.2 or Muse Voice Transcribe for audio). Tool/function calling, structured outputs, prompt caching, search grounding, and reasoning_effort (minimal/low/medium/high/xhigh; none is not supported). Compare it for Coding, Agents, Long context, and Vision.

    1.05m

    Context