Last refreshed 2026-08-26. Next refresh: weekly.
Why use Qwen3.8-Flash-Next on OpenRouter?
OpenRouter offers Qwen3.8-Flash-Next with pay-as-you-go pricing at $0.15/1M input tokens. OpenRouter is a multi-provider LLM aggregator offering unified API access to 300+ models from all major labs and emerging providers, with automatic failover for reliability.
Input / 1M
$0.15
Output / 1M
$0.47
Cache
read $0.016
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: qwen/qwen3.8-flashModel ID
qwen/qwen3.8-flashRequest example
Curated snippets for this provider are not sourced yet. Use OpenRouter documentation with model ID
qwen/qwen3.8-flash.Gotchas
- Use provider model ID "qwen/qwen3.8-flash", not the LLMReference slug "qwen3.8-flash-next".
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $0.15 |
| Output tokens | $0.47 |
Capabilities
VisionMultimodalReasoning
About Qwen3.8-Flash-Next
Qwen3.8-Flash-Next is Alibaba's experimental open-weight preview of the architecture planned for Qwen4. It is a 125B-total / 6B-active Mixture-of-Experts causal language model with a vision encoder, plus 51B n-gram embedding and 4B MTP parameters. Native context is 262,144 tokens (extensible to 1,000,000). Weights are on Hugging Face under the Qwen Community License 1.0. No first-party hosted token prices are seeded.
Get Started
Model Specs
Released2026-08-26
Parameters125B total, 6B active (+51B n-gram embedding, 4B MTP)
Context262k
ArchitectureMixture of Experts