PaddleOCR VL

Released
2025-10-16
Last refreshed
2026-06-29
Status
Researched 133d ago
Open sourceCommercial use: permittedMultimodalVision
Baidu AI releases · 9 in the last 12 months · this family litChangelog →
Specifications
Released
2025-10-16
Context
16k
Parameters
0.9B
Architecture
Encoder-Decoder
Specialization
vision
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Available
Code
Unknown
Training
Pretrained
Created by

Innovative text-to-video and app builder

Beijing, China
Founded 2010
Website
Pricing
Output / 1M
$0.020
Input / 1M
$0.020

Cheapest of 1 route · Novita AI

About

PaddleOCR VL is a 0.9B ultra-compact vision-language model from Baidu's PaddlePaddle team for multilingual document parsing. Combines a NaViT-style dynamic resolution visual encoder with ERNIE-4.5-0.3B. Supports 109 languages for recognizing text, tables, formulas, and charts. Achieved 92.56 on OmniDocBench V1.5, surpassing larger models including DeepSeek-OCR. Released October 16, 2025.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
Novita AI$0.020$0.020
Serverless

Capabilities

VisionMultimodal

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.