Kimi K3
Kimi K3 is worth evaluating for rag, agents, and long context when its provider route and context window match the workload.
Use it for
- Teams evaluating rag, agents, and long context
- Workloads that can use a 1.05m context window
- Buyers comparing 1 tracked provider route
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- Kimi
- Released
- 2026-07-14
- Context
- 1.05m
- Max output
- 1,048,576
- Parameters
- 2.8T total; 16 of 896 experts active
- Architecture
- Mixture of Experts
- Specialization
- general
- Openness
- Open weights
- License
- Kimi K3 LicenseCommercial use: conditional
- Weights
- Available
- Code
- Unknown
Cheapest of 1 route · Moonshot AI Kimi · cache read $0.300
About
Kimi K3 is Moonshot AI's 2.8-trillion-parameter flagship multimodal model for long-horizon coding, knowledge work, deep reasoning, and agentic workflows. It uses Kimi Delta Attention, Attention Residuals, and a sparse MoE design (16 of 896 experts active), supports a 1,048,576-token context window, text/image/video input, always-on reasoning, ToolCalls, strict JSON Schema structured output, automatic context caching, and partial mode through Moonshot's OpenAI-compatible API.
Top use-case fit: coding, agents, and build tasks
RAG
Included by capability and metadata signals in the decision map.
Agents
Included by capability and metadata signals in the decision map.
Long context
Included by capability and metadata signals in the decision map.
Provider price ladder
Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Cache | Route |
|---|---|---|---|---|
| Moonshot AI Kimi | $3.00 | $15.00 | read $0.300 | Serverless |
Capabilities
Benchmark peer barsfor RAG
No task-mapped benchmark peers are available for this model yet.
Benchmark scores(13)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Program Bench | 77.8 | KimiCode harness; max reasoning effortObserved 2026-07-16 | — | Source |
| Terminal-Bench 2.1 | 88.3 | Terminal-Bench 2.1; KimiCode harness; max reasoning effortObserved 2026-07-16 | — | Source |
| Google-Proof Q&A | 93.5 | GPQA-Diamond; max reasoning effortObserved 2026-07-16 | — | Source |
| Humanity's Last Exam | 43.5 | HLE-Full without tools; max reasoning effortObserved 2026-07-16 | — | Source |
| Humanity's Last Exam — With Tools | 56.0 | HLE-Full with tools; max reasoning effortObserved 2026-07-16 | — | Source |
| BrowseComp | 91.2 | BrowseComp; 300K-token context compaction; max reasoning effortObserved 2026-07-16 | — | Source |
| DeepSearchQA | 95.0 | DeepSearchQA F1; max reasoning effortObserved 2026-07-16 | — | Source |
| Toolathlon | 73.2 | Toolathlon-Verified; max reasoning effortObserved 2026-07-16 | — | Source |
| MCP-Atlas | 84.2 | 500-task public subset; 100-turn limit; Gemini 3.1 Pro judge; max reasoning effortObserved 2026-07-16 | — | Source |
| AutomationBench | 30.8 | 600-task public subset; official GitHub setup; max reasoning effortObserved 2026-07-16 | — | Source |
| WorldVQA | 51.0 | WorldVQA ForceAnswer; max reasoning effort; five-run meanObserved 2026-07-16 | — | Source |
| MMMU Pro | 81.6 | MMMU-Pro; official protocol; three-run mean; max reasoning effortObserved 2026-07-16 | — | Source |
| CharXiv | 84.8 | CharXiv RQ; three-run mean; max reasoning effortObserved 2026-07-16 | — | Source |
Migration checks
No linked migration route is available for this model yet.
Cheapest of 1 route · Moonshot AI Kimi · cache read $0.300