Qwen3.5-397B-A17B
Qwen3.5-397B-A17B is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 262k context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- Qwen3.5
- Released
- 2026-02-16
- Context
- 262k
- Parameters
- 397B
- Architecture
- Mixture of Experts
- Openness
- Open source
- License
- Apache 2.0OSI-approvedCommercial use: permitted
- Weights
- Available
- Code
- Unknown
Cheapest of 4 routes · Alibaba Cloud PAI-EAS
About
Alibaba's largest Qwen3.5 model, featuring a Mixture-of-Experts architecture with 397B total parameters and 17B active per token (using 512 total experts with 10 routed + 1 shared active). Supports 201 languages with a native 262K token context window extensible to 1M tokens via YaRN. Includes a thinking/reasoning mode, tool calling with MCP integration, and unified vision-language capabilities through early fusion training.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ C2 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ C4 relevant benchmarks in the decision map.
Provider price ladder
Compare all 4Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| Alibaba Cloud PAI-EAS | $0.390 | $2.34 | Serverless |
| OpenRouter | $0.390 | $2.34 | Serverless |
| Novita AI | $0.600 | $3.60 | Serverless |
| Together AI | $0.600 | $3.60 | Serverless |
Available via routers & gateways(1)
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(13)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Google-Proof Q&A | 89.3 | diamondObserved 2026-04-18 | — | Artificial Analysis |
| MMLU PRO | 87.8 | From official HuggingFace model card (accuracy)Observed 2026-06-07 | — | Source |
| Massive Multi-discipline Multimodal Understanding | 85.0 | —Observed 2026-04-19 | — | Source |
| Instruction-Following Evaluation | 92.6 | —Observed 2026-04-19 | — | Source |
| BFCL | 72.9 | v4Observed 2026-04-19 | — | Source |
| SWE-bench Verified | 76.2 | SWE-bench VerifiedObserved 2026-05-12 | — | Source |
| AIME 2026 | 91.3 | AIME 2026 (accuracy)Observed 2026-06-07 | — | Source |
| Berkeley Function Calling Leaderboard v3 | 72.9 | BFCL-V4, from official model card (accuracy)Observed 2026-06-07 | — | Source |
| Humanity's Last Exam | 28.7 | HLE with CoT, no tools, from official model card (accuracy)Observed 2026-06-07 | — | Source |
| LiveCodeBench | 83.6 | LiveCodeBench v6 (pass@1)Observed 2026-06-07 | — | Source |
| MultiChallenge | 67.6 | Multi-Challenge leaderboard rank 2 of 28 (accuracy%)Observed 2026-06-07 | — | Source |
| τ-bench | 86.7 | TAU2-Bench, from official model card (accuracy)Observed 2026-06-07 | — | Source |
| Terminal-Bench 2.0 | 52.5 | Terminal-Bench 2.0 (accuracy%)Observed 2026-06-07 | — | Source |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(8)
Compare Qwen3.5-397B-A17B with other models
- Qwen3.5-397B-A17B vs DeepSeek V4 Pro783
- Qwen3.5-397B-A17B vs DeepSeek V4 Flash552
- Qwen3.5-397B-A17B vs Granite Vision 4.1 4B162
- Qwen3.5-397B-A17B vs Kimi K2.5123
- Qwen3.5-397B-A17B vs Claude Sonnet 4.6111
- Qwen3.5-397B-A17B vs GLM-5.1108
- Qwen3.5-397B-A17B vs Qwen3.5-122B-A10B102
- Qwen3.5-397B-A17B vs Claude Opus 4.694
Comparison and alternatives
Browse all comparisons →Show all 53 popular comparisonssorted by 7-day search impressions
Cheapest of 4 routes · Alibaba Cloud PAI-EAS