LLM Reference
Moonshot AI Kimi

Using Kimi K2.7-Code HighSpeed on Moonshot AI Kimi

Implementation guide · Kimi K2 · Moonshot AI

ServerlessOpen Source

Moonshot AI Kimi exposes Kimi K2.7-Code HighSpeed through model ID kimi-k2.7-code-highspeed. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-22. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Moonshot AI Kimi and generate an API key.
  2. 2
    Use the Moonshot AI Kimi SDK or REST API to call kimi-k2.7-code-highspeed — see the documentation for request format.
  3. 3
    You'll be billed $1.90/1M input, $8.00/1M output tokens. See full pricing.

Code Examples

See Moonshot AI Kimi documentation for integration details.

Pricing on Moonshot AI Kimi

TypePrice (per 1M)
Input tokens$1.90
Output tokens$8.00
Image input$1.00

Capabilities

VisionMultimodalReasoningJSON / Tool useStructured OutputsPrompt Caching

About Kimi K2.7-Code HighSpeed

HighSpeed serving variant of Kimi K2.7-Code optimized for throughput at the cost of latency flexibility. Announced June 15, 2026 — three days after the standard K2.7-Code release. Delivers approximately 180 output tokens per second (up to 260 tokens/s on short-context tasks), around 6× faster than standard K2.7-Code. Same underlying 1T-parameter MoE architecture (32B active, 384 experts, 8 selected per token) with MoonViT vision encoder, 262K context window, and thinking mode always on. Best suited for interactive or latency-bound workflows; the standard variant is preferred for correctness-sensitive long-horizon agentic work.

Model Specs

Released2026-06-15
Parameters1T
Context262k
ArchitectureMixture of Experts

Provider

Moonshot AI Kimi
Moonshot AI Kimi

Moonshot AI

New York City, New York, United States