Qwen3 Models by Alibaba
Last refreshed 2026-06-01. Next refresh: weekly.
Details
Capabilities
About
Qwen3 is a family of 19 AI models by Alibaba, released in 2025.
Decision facts
- Best fit
- vision and multimodal workreasoningJSON / Tool use
- Capability starting point
- Qwen3-Max with 262k context and JSON / Tool use, structured outputs, and multimodal inputs
- Lowest tracked input
- Qwen3-8B · $0.035/1M · Novita AI
- Closest related family
- Tongyi DeepResearch
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs 128k context, 105B parameters, and JSON / Tool use.
Use when the workload needs 128k context, 8B parameters, and structured outputs.
Use when the workload needs 128k context and 8B parameters.
Use when the workload needs 128k context and 70B parameters.
Use when the workload needs 128k context and 70B parameters.
Use when the workload needs 128k context, 235B parameters, and structured outputs.
Use when the workload needs 40k context, 32B parameters, and structured outputs.
Use when the workload needs 262k context, JSON / Tool use, and structured outputs.
Use when the workload needs 128k context, 30B parameters, and structured outputs.
Use when the workload needs 256k context, 9B parameters, and structured outputs.
Use when the workload needs 160k context, 8B parameters, and reasoning.
Use when the workload needs 40k context and 600M parameters.
Use when the workload needs 40k context, 14B parameters, and structured outputs.
Use when the workload needs 40k context and 1.7B parameters.
Use when the workload needs 40k context and 4B parameters.
Use when the workload needs 160k context, 671B parameters, and reasoning.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Qwen3-105B | Use when the workload needs 128k context, 105B parameters, and JSON / Tool use. | 2025-12 | 128k context105B parametersJSON / Tool use | Current |
| Qwen3-Next-80B-A3B | Use when the workload needs structured outputs. | 2025-12 | structured outputs | Current |
| Qwen-Plus | Use when the workload needs 1m context. | 2025-11 | 1m context | Current |
| Qwen3-8B | Use when the workload needs 128k context, 8B parameters, and structured outputs. | 2025-08 | 128k context8B parametersstructured outputs | Current |
| Qwen3-8B-Instruct | Use when the workload needs 128k context and 8B parameters. | 2025-08 | 128k context8B parameters | Current |
| Qwen3-70B | Use when the workload needs 128k context and 70B parameters. | 2025-08 | 128k context70B parameters | Current |
| Qwen3-70B-Instruct | Use when the workload needs 128k context and 70B parameters. | 2025-08 | 128k context70B parameters | Current |
| Qwen-Flash | Use when the workload needs 1m context. | 2025-08 | 1m context | Current |
| Qwen3-235B-A22B | Use when the workload needs 128k context, 235B parameters, and structured outputs. | 2025-04 | 128k context235B parametersstructured outputs | Current |
| Qwen3-32B | Use when the workload needs 40k context, 32B parameters, and structured outputs. | 2025-04 | 40k context32B parametersstructured outputs | Current |
| Qwen3-Max | Use when the workload needs 262k context, JSON / Tool use, and structured outputs. | 2025-04 | 262k contextJSON / Tool usestructured outputs | Current |
| Qwen3-30B-A3B | Use when the workload needs 128k context, 30B parameters, and structured outputs. | 2025-04 | 128k context30B parametersstructured outputs | Current |
| Qwen3-9B | Use when the workload needs 256k context, 9B parameters, and structured outputs. | 2025-04 | 256k context9B parametersstructured outputs | Current |
| DeepSeek R1 0528 Distill Qwen3-8B | Use when the workload needs 160k context, 8B parameters, and reasoning. | 2025-01 | 160k context8B parametersreasoning | Current |
| Qwen3-0.6B | Use when the workload needs 40k context and 600M parameters. | 2025-01 | 40k context600M parameters | Current |
| Qwen3-14B | Use when the workload needs 40k context, 14B parameters, and structured outputs. | 2025-01 | 40k context14B parametersstructured outputs | Current |
| Qwen3-1.7B | Use when the workload needs 40k context and 1.7B parameters. | 2025-01 | 40k context1.7B parameters | Current |
| Qwen3-4B | Use when the workload needs 40k context and 4B parameters. | 2025-01 | 40k context4B parameters | Current |
| DeepSeek R1 0528 Qwen3-8B | Use when the workload needs 160k context, 671B parameters, and reasoning. | 2025-01 | 160k context671B parametersreasoning | Current |
Release Timeline
5 release groupsSpecifications(19 models)
| Model | Released | Context | Parameters | Vision | Multimodal | Reasoning | JSON / Tool use | Structured Outputs |
|---|---|---|---|---|---|---|---|---|
| Qwen3-105B | 2025-12 | 128k | 105B | No | No | No | Yes | No |
| Qwen3-Next-80B-A3B | 2025-12 | — | 80B (3B active) | No | No | No | No | Yes |
| Qwen-Plus | 2025-11 | 1m | — | No | No | No | No | No |
| Qwen3-8B | 2025-08 | 128k | 8B | No | No | No | No | Yes |
| Qwen3-8B-Instruct | 2025-08 | 128k | 8B | No | No | No | No | No |
| Qwen3-70B | 2025-08 | 128k | 70B | No | No | No | No | No |
| Qwen3-70B-Instruct | 2025-08 | 128k | 70B | No | No | No | No | No |
| Qwen-Flash | 2025-08 | 1m | — | No | No | No | No | No |
| Qwen3-235B-A22B | 2025-04 | 128k | 235B | No | No | No | No | Yes |
| Qwen3-32B | 2025-04 | 40k | 32B | No | No | No | No | Yes |
| Qwen3-Max | 2025-04 | 262k | — | Yes | Yes | No | Yes | Yes |
| Qwen3-30B-A3B | 2025-04 | 128k | 30B | No | No | No | No | Yes |
| Qwen3-9B | 2025-04 | 256k | 9B | No | No | No | No | Yes |
| DeepSeek R1 0528 Distill Qwen3-8B | 2025-01 | 160k | 8B | No | No | Yes | No | No |
| Qwen3-0.6B | 2025-01 | 40k | 0.6B | No | No | No | No | No |
| Qwen3-14B | 2025-01 | 40k | 14B | No | No | No | No | Yes |
| Qwen3-1.7B | 2025-01 | 40k | 1.7B | No | No | No | No | No |
| Qwen3-4B | 2025-01 | 40k | 4B | No | No | No | No | No |
| DeepSeek R1 0528 Qwen3-8B | 2025-01 | 160k | 671B | No | No | Yes | No | No |
Available From(11 providers)
Pricing
Popular comparisons in this family
Comparisons
- Qwen3-Max vs GPT-4o (08-06)
- Qwen3-Max vs Claude Sonnet 4.6
- Qwen3-Max vs DeepSeek V4 Pro
- Qwen3-30B-A3B vs Mistral Large 2.1 (2411)
- Qwen3-30B-A3B vs Llama 3.3 70B
- Qwen3.6-Max vs Qwen3-Max
- GPT-5 vs Qwen3-Max
- Claude Opus 4.7 vs Qwen3-Max
Models(19)
Qwen3-105B
Qwen3-Next-80B-A3B
Qwen-Plus
Qwen3-8B
Qwen3-8B-Instruct
Qwen3-70B
Qwen3-70B-Instruct
Qwen-Flash
Qwen3-235B-A22B
Qwen3-32B
Qwen3-Max
Qwen3-30B-A3B
Qwen3-9B
DeepSeek R1 0528 Distill Qwen3-8B
Qwen3-0.6B
Qwen3-14B
Qwen3-1.7B
Qwen3-4B
DeepSeek R1 0528 Qwen3-8B






