Xiaomi MiMo-V2-Flash
- Family
- MiMo V2
- Released
- 2025-12-17
- Context
- 262k
- Parameters
- 309B
- Architecture
- Mixture of Experts
- Knowledge cutoff
- 2024-12
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
- Training
- Pretrained
Cheapest of 2 routes · Novita AI
About
MiMo-V2-Flash is Xiaomi's efficient open-source Mixture-of-Experts model, announced December 17, 2025 at Xiaomi's Human-Car-Home Ecosystem Partner Conference. It has 309B total parameters with 15B active, uses hybrid attention that interleaves Sliding Window Attention and Global Attention, and extends native 32K context to 256K. Multi-Token Prediction enables about 2.6x speculative decoding speedup. The model was distributed with weights on Hugging Face and ranked highly on SWE-Bench Verified and multilingual benchmarks at research time.
Provider price ladder
Compare all 2Compare API pricing across 2 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Cache | Route |
|---|---|---|---|---|
| Novita AI | $0.100 | $0.300 | - | Serverless |
| Vercel AI Gateway | $0.100 | $0.300 | read $0.010 | Serverless |
Capabilities
Benchmark peer barsfor RAG
No task-mapped benchmark peers are available for this model yet.
Migration checks
No linked migration route is available for this model yet.
Cheapest of 2 routes · Novita AI