PaddleOCR VL Models by Baidu AI
Baidu AI releases · 9 in the last 12 months · this family litChangelog →
1 model2025Up to 16k ctxFrom $0.020/1M input
Last refreshed 2026-05-22. Next refresh: weekly.
Details
ResearcherBaidu AI
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models1
Released2025
Max context16k
Capabilities
VisionAll models
MultimodalAll models
About
PaddleOCR VL is Baidu's ultra-compact vision-language model family for multilingual document parsing and OCR. The flagship 0.9B model combines a NaViT-style dynamic resolution visual encoder with ERNIE-4.5-0.3B and supports 109 languages. Achieved SOTA on OmniDocBench V1.5 at launch.
Decision facts
- Best fit
- vision and multimodal workcoding
- Capability starting point
- PaddleOCR VL with 16k context and multimodal inputs
- Lowest tracked input
- PaddleOCR VL · $0.020/1M · Novita AI
- Closest related family
- ERNIE 4.5
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
1 in view
PaddleOCR VLCurrent
Use when the workload needs vision, 16k context, and 900M parameters.
2025-10vision16k context900M parameters
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| PaddleOCR VL | Use when the workload needs vision, 16k context, and 900M parameters. | 2025-10 | vision16k context900M parameters | Current |
Release Timeline
1 release group2025-10
1 current
PaddleOCR VL
Currentvision16k context900M parameters
Specifications(1 models)
| Model | Released | Context | Parameters | Vision | Multimodal |
|---|---|---|---|---|---|
| PaddleOCR VL | 2025-10 | 16k | 0.9B | Yes | Yes |
Available From(1 provider)
Pricing
| Model | Provider | Input / 1M | Output / 1M | Type |
|---|---|---|---|---|
| PaddleOCR VL | Novita AI | $0.02 | $0.02 | Serverless |


