LLM Reference

Qwen2.5 Models by Alibaba

AlibabaApache 2.0Open sourceHighlight
17 models2024–2025Up to 128k ctxFrom $0.03/1M input

Details

ResearcherAlibaba
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models17
Released2024–2025
Max context128k

Capabilities

Vision1 of 17 models
Multimodal1 of 17 models
Function Calling1 of 17 models
Tool Use1 of 17 models
Structured Outputs4 of 17 models

About

The Qwen 2.5 large language model (LLM) family, developed by Alibaba Cloud's Qwen team, consists of decoder-only dense models that are open-sourced and come in seven different sizes ranging from 0.5 billion to 72 billion parameters 124. Built on a colossal dataset of up to 18 trillion tokens, these models showcase improvements from the Qwen 2 series, especially in knowledge, coding, and mathematical tasks 13. They boast enhanced capabilities in instruction following, long-text generation, and understanding structured data, including generating structured outputs like JSON 23. Other noteworthy features include improved system prompt handling for better role-playing and chatbot configuration 23. The models support over 29 languages, including Chinese and English, and include specialized versions like Qwen2.5-Coder and Qwen2.5-Math for specific tasks 236. Moreover, Qwen-Plus and Qwen-Turbo are accessible through Alibaba Cloud Model Studio APIs 3.

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

17 in view

Use when the workload needs 128k context, 72B parameters, and tool use.

2025-10128k context72B parameterstool use

Use when the workload needs 32k context.

2025-0132k context

Use when the workload needs 33k context, 72B parameters, and multimodal inputs.

2025-0133k context72B parametersmultimodal inputs

Use when the workload needs 128k context and 490M parameters.

2024-06128k context490M parameters

Use when the workload needs 128k context and 490M parameters.

2024-06128k context490M parameters

Use when the workload needs 128k context and 1.5B parameters.

2024-06128k context1.5B parameters

Use when the workload needs 128k context and 1.5B parameters.

2024-06128k context1.5B parameters
Qwen2.5-3BCurrent

Use when the workload needs 128k context and 3.1B parameters.

2024-06128k context3.1B parameters

Use when the workload needs 128k context and 3.1B parameters.

2024-06128k context3.1B parameters
Qwen2.5-7BCurrent

Use when the workload needs 128k context and 7.6B parameters.

2024-06128k context7.6B parameters

Use when the workload needs 128k context, 7.6B parameters, and structured outputs.

2024-06128k context7.6B parametersstructured outputs

Use when the workload needs 128k context and 14.7B parameters.

2024-06128k context14.7B parameters

Use when the workload needs 128k context, 14.7B parameters, and structured outputs.

2024-06128k context14.7B parametersstructured outputs

Use when the workload needs 128k context and 32.5B parameters.

2024-06128k context32.5B parameters

Use when the workload needs 128k context, 32.5B parameters, and structured outputs.

2024-06128k context32.5B parametersstructured outputs

Use when the workload needs 128k context and 72.7B parameters.

2024-06128k context72.7B parameters

Use when the workload needs 128k context, 72.7B parameters, and structured outputs.

2024-06128k context72.7B parametersstructured outputs

Release Timeline

3 release groups
2025-10
1 current
Qwen2.5-72B
128k context72B parameterstool use
Current
2025-01
2 current
Qwen2.5-Max
32k context
Current
Qwen2.5-VL-72B
33k context72B parametersmultimodal inputs
Current
2024-06
14 current
Qwen2.5-0.5B
128k context490M parameters
Current
Qwen2.5-0.5B-Instruct
128k context490M parameters
Current
Qwen2.5-1.5B
128k context1.5B parameters
Current
Qwen2.5-1.5B-Instruct
128k context1.5B parameters
Current
Qwen2.5-14B
128k context14.7B parameters
Current
Qwen2.5-14B-Instruct
128k context14.7B parametersstructured outputs
Current
Qwen2.5-32B
128k context32.5B parameters
Current
Qwen2.5-32B-Instruct
128k context32.5B parametersstructured outputs
Current
Qwen2.5-3B
128k context3.1B parameters
Current
Qwen2.5-3B-Instruct
128k context3.1B parameters
Current
Qwen2.5-72B
128k context72.7B parameters
Current
Qwen2.5-72B-Instruct
128k context72.7B parametersstructured outputs
Current
Qwen2.5-7B
128k context7.6B parameters
Current
Qwen2.5-7B-Instruct
128k context7.6B parametersstructured outputs
Current

Specifications(17 models)

Qwen2.5 model specifications comparison
ModelReleasedContextParametersVisionMultimodalFn CallingTool UseStructured Outputs
Qwen2.5-72B2025-10128k72BNoNoYesYesNo
Qwen2.5-Max2025-0132kNoNoNoNoNo
Qwen2.5-VL-72B2025-0133k72BYesYesNoNoNo
Qwen2.5-0.5B2024-06128k490MNoNoNoNoNo
Qwen2.5-0.5B-Instruct2024-06128k490MNoNoNoNoNo
Qwen2.5-1.5B2024-06128k1.54BNoNoNoNoNo
Qwen2.5-1.5B-Instruct2024-06128k1.54BNoNoNoNoNo
Qwen2.5-3B2024-06128k3.09BNoNoNoNoNo
Qwen2.5-3B-Instruct2024-06128k3.09BNoNoNoNoNo
Qwen2.5-7B2024-06128k7.61BNoNoNoNoNo
Qwen2.5-7B-Instruct2024-06128k7.61BNoNoNoNoYes
Qwen2.5-14B2024-06128k14.7BNoNoNoNoNo
Qwen2.5-14B-Instruct2024-06128k14.7BNoNoNoNoYes
Qwen2.5-32B2024-06128k32.5BNoNoNoNoNo
Qwen2.5-32B-Instruct2024-06128k32.5BNoNoNoNoYes
Qwen2.5-72B2024-06128k72.7BNoNoNoNoNo
Qwen2.5-72B-Instruct2024-06128k72.7BNoNoNoNoYes

Pricing

Qwen2.5 model pricing by provider
ModelProviderInput / 1MOutput / 1MType
Qwen2.5-7B-InstructDeepInfra$0.03$0.03Serverless
Qwen2.5-7B-InstructOpenRouter$0.04$0.1Serverless
Qwen2.5-7B-InstructSiliconFlow$0.04$0.04Serverless
Qwen2.5-7B-InstructNovita AI$0.07$0.07Serverless
Qwen2.5-14B-InstructSiliconFlow$0.08$0.08Serverless
Qwen2.5-14B-InstructDeepInfra$0.1$0.1Serverless
Qwen2.5-0.5B-InstructFireworks AI$0.1$0.1Serverless
Qwen2.5-1.5B-InstructFireworks AI$0.1$0.1Serverless
Qwen2.5-0.5BBitdeer AI$0.12$0.36Serverless
Qwen2.5-1.5BBitdeer AI$0.12$0.36Serverless
Qwen2.5-7BBitdeer AI$0.12$0.36Serverless
Qwen2.5-14BBitdeer AI$0.12$0.36Serverless
Qwen2.5-32BBitdeer AI$0.12$0.36Serverless
Qwen2.5-7B-InstructTogether AI$0.15$0.15Serverless
Qwen2.5-32B-InstructSiliconFlow$0.15$0.15Serverless
Qwen2.5-72B-InstructChutes AI$0.18$0.54Serverless
Qwen2.5-14B-InstructFireworks AI$0.2$0.2Serverless
Qwen2.5-7BFireworks AI$0.2$0.2Serverless
Qwen2.5-14BFireworks AI$0.2$0.2Serverless
Qwen2.5-7B-InstructFireworks AI$0.2$0.2Serverless
Qwen2.5-72BBitdeer AI$0.2$0.6Serverless
Qwen2.5-72B-InstructSiliconFlow$0.28$0.28Serverless
Qwen2.5-72B-InstructDeepInfra$0.36$0.4Serverless
Qwen2.5-72B-InstructOpenRouter$0.36$0.4Serverless
Qwen2.5-72B-InstructNovita AI$0.38$0.4Serverless
Qwen2.5-32B-InstructReplicate API$0.6$0.6Serverless
Qwen2.5-VL-72BNovita AI$0.8$0.8Serverless
Qwen2.5-32BFireworks AI$0.9$0.9Serverless
Qwen2.5-32B-InstructFireworks AI$0.9$0.9Serverless
Qwen2.5-72BFireworks AI$0.9$0.9Serverless
Qwen2.5-72B-InstructFireworks AI$0.9$0.9Serverless
Qwen2.5-72B-InstructReplicate API$1.3$1.3Serverless

Popular comparisons in this family

Frequently Asked Questions

What is Qwen2.5 used for?
Qwen2.5 is used for vision and multimodal work, agent workflows and tool use, and structured outputs. The family description and listed model capabilities point to those workloads as the best fit.
How does Qwen2.5 compare to Tongyi DeepResearch?
Qwen2.5 by Alibaba is strongest where you need vision and multimodal work, while Tongyi DeepResearch by Alibaba is the closest related family to check for adjacent model selection. Qwen2.5 has 17 listed variants and reaches up to 128k context, while Tongyi DeepResearch reaches up to 131k context, so compare the specs and pricing tables before choosing a production model.
Which Qwen2.5 model should I use?
For the lowest listed input price, start with Qwen2.5-7B-Instruct through DeepInfra at $0.03/1M input tokens. For the most capable/latest local choice, evaluate Qwen2.5-72B with 128k context and tool use and function calling.

Models(17)