Qwen3-VL Models by Alibaba

AlibabaApache 2.0Open source
Alibaba releases · 46 in the last 12 months · this family litChangelog →
6 models2025Up to 256k ctxFrom $0.080/1M input

Last refreshed 2026-06-04. Next refresh: weekly.

Details

ResearcherAlibaba
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models6
Released2025
Max context256k

Capabilities

VisionAll models
MultimodalAll models
JSON / Tool use4 of 6 models
Structured OutputsAll models

About

Qwen3-VL is Alibaba’s Qwen Team vision–language series for interleaved text, image, and video understanding. The family ships dense sizes (2B/4B/8B/32B) and MoE sizes (30B-A3B and flagship 235B-A22B), each with Instruct and Thinking post-trained editions and a native long multimodal context window. Qwen documents stronger visual agents (GUI operation), spatial grounding, OCR, STEM reasoning, and long video/document work, with architecture updates such as Interleaved-MRoPE, DeepStack vision fusion, and text-timestamp video alignment.

Decision facts

Best fit
vision and multimodal workJSON / Tool usestructured outputs
Capability starting point
Qwen3-VL-235B-A22B with 128k context and structured outputs and multimodal inputs
Lowest tracked input
Qwen3 VL 8B Instruct · $0.080/1M · Novita AI
Closest related family
Ovis Image

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

6 in view

Use when the workload needs 128k context, structured outputs, and multimodal inputs.

2025-12128k contextstructured outputsmultimodal inputs

Use when the workload needs 256k context, 235B parameters, and JSON / Tool use.

2025-09256k context235B parametersJSON / Tool use

Use when the workload needs 128k context, 30B parameters, and JSON / Tool use.

2025-09128k context30B parametersJSON / Tool use

Use when the workload needs 128k context, 32B parameters, and JSON / Tool use.

2025-09128k context32B parametersJSON / Tool use

Use when the workload needs 128k context, 8B parameters, and JSON / Tool use.

2025-09128k context8B parametersJSON / Tool use

Use when the workload needs multimodal, 128k context, and 30B parameters.

2025-01multimodal128k context30B parameters

Release Timeline

3 release groups
2025-12
1 current
Qwen3-VL-235B-A22B
128k contextstructured outputsmultimodal inputs
Current
2025-09
4 current
Qwen3 VL 235B A22B Instruct
256k context235B parametersJSON / Tool use
Current
Qwen3 VL 30B A3B Instruct
128k context30B parametersJSON / Tool use
Current
Qwen3 VL 32B Instruct
128k context32B parametersJSON / Tool use
Current
Qwen3 VL 8B Instruct
128k context8B parametersJSON / Tool use
Current
2025-01
1 current
Qwen3-VL-30B-A3B
multimodal128k context30B parameters
Current

Specifications(6 models)

Qwen3-VL model specifications comparison
ModelReleasedContextParametersVisionMultimodalJSON / Tool useStructured Outputs
Qwen3-VL-235B-A22B2025-12128k235B (22B active)YesYesNoYes
Qwen3 VL 235B A22B Instruct2025-09256k235BYesYesYesYes
Qwen3 VL 30B A3B Instruct2025-09128k30BYesYesYesYes
Qwen3 VL 32B Instruct2025-09128k32BYesYesYesYes
Qwen3 VL 8B Instruct2025-09128k8BYesYesYesYes
Qwen3-VL-30B-A3B2025-01128k30BYesYesNoYes

Pricing

Qwen3-VL model pricing by provider
ModelProviderInput / 1MOutput / 1MType
Qwen3 VL 8B InstructNovita AI$0.08$0.5Serverless
Qwen3 VL 32B InstructOpenRouter$0.104$0.416Serverless
Qwen3 VL 8B InstructOpenRouter$0.117$0.455Serverless
Qwen3-VL-30B-A3BOpenRouter$0.13$1.56Serverless
Qwen3 VL 30B A3B InstructOpenRouter$0.13$0.52Serverless
Qwen3-VL-30B-A3BFireworks AI$0.15$0.6Serverless
Qwen3 VL 30B A3B InstructNovita AI$0.2$0.7Serverless
Qwen3 VL 235B A22B InstructOpenRouter$0.21$1.9Serverless
Qwen3-VL-235B-A22BOpenRouter$0.26$2.6Serverless
Qwen3 VL 235B A22B InstructNovita AI$0.3$1.5Serverless
Qwen3 VL 235B A22B InstructVercel AI Gateway$0.4$1.6Serverless
Qwen3-VL-235B-A22BAWS Bedrock$0.53$2.66Serverless
Qwen3-VL-235B-A22BNovita AI$0.98$3.95Serverless