Qwen2-VL-72B-Instruct

Released
2025-01-01
Last refreshed
2026-05-19
Status
Researched 134d ago
Open sourceCommercial use: permittedMultimodalVision
Alibaba releases · 46 in the last 12 monthsChangelog →
Specifications
Family
Qwen2-VL
Released
2025-01-01
Context
32k
Parameters
72B
Architecture
Decoder Only
Knowledge cutoff
2023-06
Specialization
multimodal
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Unknown
Code
Unknown
Training
Pretrained
Created by

AI research institute of Alibaba Group.

Hangzhou, Zhejiang, China
Founded 2017
Website
Pricing
Output / 1M
$0.900
Input / 1M
$0.900

Cheapest of 1 route · Fireworks AI

About

Qwen2-VL-72B-Instruct is Alibaba's Qwen2-VL model focused on multimodal input across text, image, and beyond. It offers a 32K-token context window.

Top use-case fit

Vision

Q/$ B

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
Fireworks AI$0.900$0.900
Serverless

Available via routers & gateways(1)

Capabilities

VisionMultimodal

Benchmark peer barsfor Vision

Benchmark scores(2)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
Massive Multi-discipline Multimodal Understanding64.5—Observed 2026-04-15—Source
MMMU Pro59.3standard 4-option (original paper harness)Observed 2024-09-04—Source

Migration checks

No linked migration route is available for this model yet.

Compare Qwen2-VL-72B-Instruct with other models