LLM Reference
OpenRouter

Using Qwen3.8-Flash-Next on OpenRouter

Implementation guide · Qwen3.8 · Alibaba

ServerlessOpen Weights

OpenRouter exposes Qwen3.8-Flash-Next through model ID qwen/qwen3.8-flash. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-08-26. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at OpenRouter and generate an API key.
  2. 2
    Use the OpenRouter SDK or REST API to call qwen/qwen3.8-flash — see the documentation for request format.
  3. 3
    You'll be billed $0.15/1M input, $0.47/1M output tokens. See full pricing.

Code Examples

See OpenRouter documentation for integration details.

Pricing on OpenRouter

TypePrice (per 1M)
Input tokens$0.15
Output tokens$0.47

Capabilities

VisionMultimodalReasoning

About Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is Alibaba's experimental open-weight preview of the architecture planned for Qwen4. It is a 125B-total / 6B-active Mixture-of-Experts causal language model with a vision encoder, plus 51B n-gram embedding and 4B MTP parameters. Native context is 262,144 tokens (extensible to 1,000,000). Weights are on Hugging Face under the Qwen Community License 1.0. No first-party hosted token prices are seeded.

Model Specs

Released2026-08-26
Parameters125B total, 6B active (+51B n-gram embedding, 4B MTP)
Context262k
ArchitectureMixture of Experts

Provider

OpenRouter
OpenRouter

OpenRouter, Inc.

New York, NY, USA