Qwen2.5 Models by Alibaba
Details
Capabilities
About
The Qwen 2.5 large language model (LLM) family, developed by Alibaba Cloud's Qwen team, consists of decoder-only dense models that are open-sourced and come in seven different sizes ranging from 0.5 billion to 72 billion parameters 124. Built on a colossal dataset of up to 18 trillion tokens, these models showcase improvements from the Qwen 2 series, especially in knowledge, coding, and mathematical tasks 13. They boast enhanced capabilities in instruction following, long-text generation, and understanding structured data, including generating structured outputs like JSON 23. Other noteworthy features include improved system prompt handling for better role-playing and chatbot configuration 23. The models support over 29 languages, including Chinese and English, and include specialized versions like Qwen2.5-Coder and Qwen2.5-Math for specific tasks 236. Moreover, Qwen-Plus and Qwen-Turbo are accessible through Alibaba Cloud Model Studio APIs 3.
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs 128k context, 72B parameters, and tool use.
Use when the workload needs 33k context, 72B parameters, and multimodal inputs.
Use when the workload needs 128k context and 490M parameters.
Use when the workload needs 128k context and 490M parameters.
Use when the workload needs 128k context and 1.5B parameters.
Use when the workload needs 128k context and 1.5B parameters.
Use when the workload needs 128k context and 3.1B parameters.
Use when the workload needs 128k context and 3.1B parameters.
Use when the workload needs 128k context and 7.6B parameters.
Use when the workload needs 128k context, 7.6B parameters, and structured outputs.
Use when the workload needs 128k context and 14.7B parameters.
Use when the workload needs 128k context, 14.7B parameters, and structured outputs.
Use when the workload needs 128k context and 32.5B parameters.
Use when the workload needs 128k context, 32.5B parameters, and structured outputs.
Use when the workload needs 128k context and 72.7B parameters.
Use when the workload needs 128k context, 72.7B parameters, and structured outputs.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Qwen2.5-72B | Use when the workload needs 128k context, 72B parameters, and tool use. | 2025-10 | 128k context72B parameterstool use | Current |
| Qwen2.5-Max | Use when the workload needs 32k context. | 2025-01 | 32k context | Current |
| Qwen2.5-VL-72B | Use when the workload needs 33k context, 72B parameters, and multimodal inputs. | 2025-01 | 33k context72B parametersmultimodal inputs | Current |
| Qwen2.5-0.5B | Use when the workload needs 128k context and 490M parameters. | 2024-06 | 128k context490M parameters | Current |
| Qwen2.5-0.5B-Instruct | Use when the workload needs 128k context and 490M parameters. | 2024-06 | 128k context490M parameters | Current |
| Qwen2.5-1.5B | Use when the workload needs 128k context and 1.5B parameters. | 2024-06 | 128k context1.5B parameters | Current |
| Qwen2.5-1.5B-Instruct | Use when the workload needs 128k context and 1.5B parameters. | 2024-06 | 128k context1.5B parameters | Current |
| Qwen2.5-3B | Use when the workload needs 128k context and 3.1B parameters. | 2024-06 | 128k context3.1B parameters | Current |
| Qwen2.5-3B-Instruct | Use when the workload needs 128k context and 3.1B parameters. | 2024-06 | 128k context3.1B parameters | Current |
| Qwen2.5-7B | Use when the workload needs 128k context and 7.6B parameters. | 2024-06 | 128k context7.6B parameters | Current |
| Qwen2.5-7B-Instruct | Use when the workload needs 128k context, 7.6B parameters, and structured outputs. | 2024-06 | 128k context7.6B parametersstructured outputs | Current |
| Qwen2.5-14B | Use when the workload needs 128k context and 14.7B parameters. | 2024-06 | 128k context14.7B parameters | Current |
| Qwen2.5-14B-Instruct | Use when the workload needs 128k context, 14.7B parameters, and structured outputs. | 2024-06 | 128k context14.7B parametersstructured outputs | Current |
| Qwen2.5-32B | Use when the workload needs 128k context and 32.5B parameters. | 2024-06 | 128k context32.5B parameters | Current |
| Qwen2.5-32B-Instruct | Use when the workload needs 128k context, 32.5B parameters, and structured outputs. | 2024-06 | 128k context32.5B parametersstructured outputs | Current |
| Qwen2.5-72B | Use when the workload needs 128k context and 72.7B parameters. | 2024-06 | 128k context72.7B parameters | Current |
| Qwen2.5-72B-Instruct | Use when the workload needs 128k context, 72.7B parameters, and structured outputs. | 2024-06 | 128k context72.7B parametersstructured outputs | Current |
Release Timeline
3 release groupsSpecifications(17 models)
| Model | Released | Context | Parameters | Vision | Multimodal | Fn Calling | Tool Use | Structured Outputs |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-72B | 2025-10 | 128k | 72B | No | No | Yes | Yes | No |
| Qwen2.5-Max | 2025-01 | 32k | — | No | No | No | No | No |
| Qwen2.5-VL-72B | 2025-01 | 33k | 72B | Yes | Yes | No | No | No |
| Qwen2.5-0.5B | 2024-06 | 128k | 490M | No | No | No | No | No |
| Qwen2.5-0.5B-Instruct | 2024-06 | 128k | 490M | No | No | No | No | No |
| Qwen2.5-1.5B | 2024-06 | 128k | 1.54B | No | No | No | No | No |
| Qwen2.5-1.5B-Instruct | 2024-06 | 128k | 1.54B | No | No | No | No | No |
| Qwen2.5-3B | 2024-06 | 128k | 3.09B | No | No | No | No | No |
| Qwen2.5-3B-Instruct | 2024-06 | 128k | 3.09B | No | No | No | No | No |
| Qwen2.5-7B | 2024-06 | 128k | 7.61B | No | No | No | No | No |
| Qwen2.5-7B-Instruct | 2024-06 | 128k | 7.61B | No | No | No | No | Yes |
| Qwen2.5-14B | 2024-06 | 128k | 14.7B | No | No | No | No | No |
| Qwen2.5-14B-Instruct | 2024-06 | 128k | 14.7B | No | No | No | No | Yes |
| Qwen2.5-32B | 2024-06 | 128k | 32.5B | No | No | No | No | No |
| Qwen2.5-32B-Instruct | 2024-06 | 128k | 32.5B | No | No | No | No | Yes |
| Qwen2.5-72B | 2024-06 | 128k | 72.7B | No | No | No | No | No |
| Qwen2.5-72B-Instruct | 2024-06 | 128k | 72.7B | No | No | No | No | Yes |
Available From(10 providers)
Pricing
Popular comparisons in this family
- Llama 3.3 70B vs Qwen2.5-72B1K
- Llama 3.3 70B Instruct (free) vs Qwen2.5-72B-Instruct1K
- Llama 3.3 70B vs Qwen2.5-72B-Instruct859
- Qwen2.5-72B-Instruct vs Qwen3.6-27B706
- DeepSeek V3 vs Qwen2.5-72B608
- Llama 3.1 405B vs Qwen2.5-72B535
- DeepSeek R1 vs Qwen2.5-72B-Instruct522
- Llama 3 70B Instruct vs Qwen2.5-72B-Instruct461
- DeepSeek V4 Flash vs Qwen2.5-72B-Instruct289
- DeepSeek V4 Pro vs Qwen2.5-72B-Instruct221
Comparisons
All comparisons →Frequently Asked Questions
- What is Qwen2.5 used for?
- Qwen2.5 is used for vision and multimodal work, agent workflows and tool use, and structured outputs. The family description and listed model capabilities point to those workloads as the best fit.
- How does Qwen2.5 compare to Tongyi DeepResearch?
- Qwen2.5 by Alibaba is strongest where you need vision and multimodal work, while Tongyi DeepResearch by Alibaba is the closest related family to check for adjacent model selection. Qwen2.5 has 17 listed variants and reaches up to 128k context, while Tongyi DeepResearch reaches up to 131k context, so compare the specs and pricing tables before choosing a production model.
- Which Qwen2.5 model should I use?
- For the lowest listed input price, start with Qwen2.5-7B-Instruct through DeepInfra at $0.03/1M input tokens. For the most capable/latest local choice, evaluate Qwen2.5-72B with 128k context and tool use and function calling.
Models(17)
Qwen2.5-72B
Qwen2.5-Max
Qwen2.5-VL-72B
Qwen2.5-0.5B
Qwen2.5-0.5B-Instruct
Qwen2.5-1.5B
Qwen2.5-1.5B-Instruct
Qwen2.5-3B
Qwen2.5-3B-Instruct
Qwen2.5-7B
Qwen2.5-7B-Instruct
Qwen2.5-14B
Qwen2.5-14B-Instruct
Qwen2.5-32B
Qwen2.5-32B-Instruct
Qwen2.5-72B
Qwen2.5-72B-Instruct






