Phi 3.5 Vision Instruct
Phi 3.5 Vision Instruct is a released long context and vision model with open-source and 128k context; evaluate it while provider pricing coverage matures.
Use it for
- Teams evaluating long context and vision
- Workloads that can use a 128k context window
Do not use it for
- Cost-sensitive launches that need sourced token pricing
- Strict JSON or tool-calling flows
- Teams that need a tracked hosted API route today
Advancing the state-of-the-art in AI and computing.
No tracked provider token pricing is available yet.
About
Phi 3.5 Vision Instruct is Microsoft Research's Phi-3 model with multimodal text and image input. It offers a 128K-token context window with weights openly available for self-hosting and scores 43 on MMMU.
Top use-case fit
Long context
Included by capability and metadata signals in the decision map.
Vision
1 relevant benchmark in the decision map.
Provider price ladder
No tracked provider token pricing is available for this model yet.
Capabilities
Benchmark peer barsfor Vision
Benchmark scores(1)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Massive Multi-discipline Multimodal Understanding | 43.0 | —Observed 2026-04-15 | — | Source |
Migration checks
No linked migration route is available for this model yet.
Compare Phi 3.5 Vision Instruct with other models
Advancing the state-of-the-art in AI and computing.
No tracked provider token pricing is available yet.