Arrow 2 Telos
- Context
- 1.05m
- Output (from)
- $30.00 / 1M
Last refreshed 2026-09-21. Next refresh: weekly.
Best LLMs for RAG in 2026 — ranked by context window, retrieval benchmarks, and tool support. Covers document QA to enterprise search.
Verdict
GPT-6 Astra is the runner-up: 1.05m vs 1.05m on Context.
RAG picks emphasize the strongest sourced long-document benchmark among tracked RAG needles, QA, and retrieval suites, then context window, then recency.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Nemotron 3 Super-120B-A12B Signal used: RULER 96.33% | $0.09 | $0.45 | |
| 2 | Llama 4 Scout 17B-16E Instruct Vision Signal used: Context 10m | $0.08 | $0.22 | |
| 3 | Gemini 1.5 Pro Signal used: Context 2m | $1.25 | $5.00 | |
| 4 | Arrow 2 Telos VisionTools Signal used: Context 1.05m | $6.00 | $30.00 | |
| 5 | GPT-6 Astra RestrictedReasoningVisionTools Signal used: Context 1.05m | $10.00 | $50.00 | |
| 6 | GPT-5.6 Luna ReasoningVisionTools Signal used: Context 1.05m | $1.00 | $6.00 | |
| 7 | GPT-5.6 Sol ReasoningVisionTools Signal used: Context 1.05m | $5.00 | $30.00 | |
| 8 | GPT-5.6 Terra ReasoningVisionTools Signal used: Context 1.05m | $2.50 | $15.00 | |
| 9 | GPT-5.5 ReasoningVisionTools Signal used: Context 1.05m | $5.00 | $30.00 | |
| 10 | GPT-5.5 Pro ReasoningVisionTools Signal used: Context 1.05m | $30.00 | $180.00 | |
| 11 | GPT-5.4 ReasoningVisionTools Signal used: Context 1.05m | $2.50 | $15.00 | |
| 12 | GPT-5.4 Pro ReasoningVisionTools Signal used: Context 1.05m | $30.00 | $180.00 | |
| 13 | Xiaomi MiMo-V2.6-Pro ReasoningVisionTools Signal used: Context 1.05m | $0.43 | $0.87 | |
| 14 | Nex-N2.5-Max ReasoningTools Signal used: Context 1.05m | — | — | |
| 15 | Gemini 3.8 Flash ReasoningVisionTools Signal used: Context 1.05m | $0.75 | $3.75 | |
| 16 | Muse Spark 1.3 ReasoningVisionTools Signal used: Context 1.05m | $1.25 | $4.25 | |
| 17 | Muse Spark 1.3 Contributor ReasoningVisionTools Signal used: Context 1.05m | $0.10 | $0.20 | |
| 18 | Spark X2.5 1.7B ReasoningTools Signal used: Context 1.05m | — | — | |
| 19 | Spark X2.5 4B ReasoningTools Signal used: Context 1.05m | — | — | |
| 20 | Hunyuan Hy4 Preview PreviewReasoningTools Signal used: Context 1.05m | $0.83 | $2.50 |
Nex-N2.5-Max is Nex AGI's trillion-scale text-only MoE agentic model in the Nex-N2.5 family, released September 8, 2026. Apache-2 weights on Hugging Face (nex-agi/Nex-N2.5-Max). 1.6T total / 49B active MoE with a native 1,048,576-token context window (config max_position_embeddings). Supports reasoning, function calling, and tool use; no vision or multimodal inputs. No OpenRouter listing in this seed pass. Sibling Nex-N2.5-mini is the multimodal MoE variant; Nex-N2.5-Pro is held from this seed pass.
1.05m
Context
Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8. Compare it for Coding, Agents, Long context, Vision, and JSON / Tool use.
1.05m
Context
Muse Spark 1.3 is Meta Superintelligence Labs' latest Muse Spark checkpoint, released September 2, 2026. Official Meta Model API model ID: muse-spark-1.3. First-party models docs: the latest version, recommended for new work, tuned for agentic workflows (multi-step tool, browser, and long-horizon tasks) with improved coding over 1.2; default model in docs examples. Product line (dev.meta.ai): Built for long-horizon coding and agentic work. Cleaner code, fewer tokens. First-party research post: rolling out today in Muse Code and Meta Model API; previously available reasoning modes available today with max reasoning coming shortly after additional safety testing. Official table: text, image, video, audio*, PDF input and text output; 1,048,576-token context. Official quickstart limits: context 1048576, output 131072; input modalities text/image/pdf/video, output text. Audio understanding in 1.3 is currently not fully supported (use 1.2 or Muse Voice Transcribe for audio). Tool/function calling, structured outputs, prompt caching, search grounding, and reasoning_effort (minimal/low/medium/high/xhigh; none is not supported). Compare it for Coding, Agents, Long context, and Vision.
1.05m
Context