LLM Reference

Best LLM for translation

Last refreshed 2026-09-21. Next refresh: weekly.

Compare multilingual LLMs and dedicated translation models for text, document, and live speech translation. Ranked by context length and general-language benchmarks until translation-specific leaderboard rows land in seed data.

Verdict

Use Arrow 2 Telos for this task today.

GPT-6 Astra is the runner-up: Current leader vs No. 2 on Pick.

Researched 4d agoWhy this pickMethodology

How we rank

Translation picks are methodology-forward until dedicated translation benchmark rows land in seed data. We surface tagged translation models, Qwen-MT specialists, and long-context chat models teams commonly route for multilingual work.

  1. Eligibility — Models with translation use-case tags, translation specialization, Qwen-MT family membership, or ≥128K context on general chat-completion routes (excluding embeddings and media-only specialists).
  2. Primary ranking — Declared context window (wider first), then MMLU when present, then newer release. Realtime speech translation models appear in the table but are labeled separately from text-generation LLMs.
  3. Benchmark gap — No translation-quality leaderboard is standardized in seed yet — do not treat this page as a definitive translation-quality ranking until WMT or vendor-neutral multilingual rows are sourced.
  4. Pricing column — Lowest tracked provider input/output where a public rate card exists.
#ModelInput $/1MOutput $/1M
1RWKV-7 Goose 0.1B
——
2RWKV-7 Goose 0.4B
——
3RWKV-7 Goose 1.5B
——
4RWKV-7 Goose 2.9B
——
5RWKV-6 Finch 14B
——
6RWKV-6 Finch 1.6B
——
7RWKV-6 Finch 3B
——
8RWKV-6 Finch 7B
——
9LTM-2-mini
——
10Llama 4 Scout 17B-16E Instruct
Vision
$0.08$0.22
11LTM-1
——
12MiniMax-01
Vision
$0.20$1.10
13Gemini 3.5 Pro
PreviewReasoningVision
——
14Gemini 1.5 Pro 002
——
15Gemini 1.5 Pro Experimental 0827
——
16Gemini 1.5 Pro Experimental 0801
——
17Gemini 1.5 Pro
$1.25$5.00
18GPT-5.5
ReasoningVisionTools
$5.00$30.00
19Arrow 2 Telos
VisionTools
$6.00$30.00
20GPT-6 Astra
RestrictedReasoningVisionTools
$10.00$50.00

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Nex-N2.5-Max is Nex AGI's trillion-scale text-only MoE agentic model in the Nex-N2.5 family, released September 8, 2026. Apache-2 weights on Hugging Face (nex-agi/Nex-N2.5-Max). 1.6T total / 49B active MoE with a native 1,048,576-token context window (config max_position_embeddings). Supports reasoning, function calling, and tool use; no vision or multimodal inputs. No OpenRouter listing in this seed pass. Sibling Nex-N2.5-mini is the multimodal MoE variant; Nex-N2.5-Pro is held from this seed pass.

    See leaderboard

    Rank

  • Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8. Compare it for Coding, Agents, Long context, Vision, and JSON / Tool use.

    See leaderboard

    Rank

  • Muse Spark 1.3 is Meta Superintelligence Labs' latest Muse Spark checkpoint, released September 2, 2026. Official Meta Model API model ID: muse-spark-1.3. First-party models docs: the latest version, recommended for new work, tuned for agentic workflows (multi-step tool, browser, and long-horizon tasks) with improved coding over 1.2; default model in docs examples. Product line (dev.meta.ai): Built for long-horizon coding and agentic work. Cleaner code, fewer tokens. First-party research post: rolling out today in Muse Code and Meta Model API; previously available reasoning modes available today with max reasoning coming shortly after additional safety testing. Official table: text, image, video, audio*, PDF input and text output; 1,048,576-token context. Official quickstart limits: context 1048576, output 131072; input modalities text/image/pdf/video, output text. Audio understanding in 1.3 is currently not fully supported (use 1.2 or Muse Voice Transcribe for audio). Tool/function calling, structured outputs, prompt caching, search grounding, and reasoning_effort (minimal/low/medium/high/xhigh; none is not supported). Compare it for Coding, Agents, Long context, and Vision.

    See leaderboard

    Rank