PaddleOCR VL
Released
2025-10-16
Last refreshed
2026-06-29
Status
Researched 133d ago
Open sourceCommercial use: permittedMultimodalVision
Baidu AI releases · 9 in the last 12 months · this family litChangelog →
Specifications
- Family
- PaddleOCR VL
- Released
- 2025-10-16
- Context
- 16k
- Parameters
- 0.9B
- Architecture
- Encoder-Decoder
- Specialization
- vision
- Openness
- Open source
- License
- Apache 2.0OSI-approvedCommercial use: permitted
- Weights
- Available
- Code
- Unknown
- Training
- Pretrained
Pricing
Output / 1M
$0.020
Input / 1M
$0.020
Cheapest of 1 route · Novita AI
Links
About
PaddleOCR VL is a 0.9B ultra-compact vision-language model from Baidu's PaddlePaddle team for multilingual document parsing. Combines a NaViT-style dynamic resolution visual encoder with ERNIE-4.5-0.3B. Supports 109 languages for recognizing text, tables, formulas, and charts. Achieved 92.56 on OmniDocBench V1.5, surpassing larger models including DeepSeek-OCR. Released October 16, 2025.
Provider price ladder
Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| Novita AI | $0.020 | $0.020 | Serverless |
Capabilities
VisionMultimodal
Benchmark peer barsfor Vision
No task-mapped benchmark peers are available for this model yet.
Migration checks
No linked migration route is available for this model yet.
Pricing
Output / 1M
$0.020
Input / 1M
$0.020
Cheapest of 1 route · Novita AI
Links