Last refreshed 2026-06-29. Next refresh: weekly.
Why use PaddleOCR VL on Novita AI?
Novita AI offers PaddleOCR VL with pay-as-you-go pricing at $0.02/1M input tokens. Novita AI offers a GPU-based inference API for image, video, and language model generation with a broad catalog of open-source models.
Input / 1M
$0.020
Output / 1M
$0.020
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: paddleocr-vlModel ID
paddleocr-vlRequest example
Curated snippets for this provider have not been sourced yet.
Gotchas
No curated gotchas have been sourced for this exact provider/model route yet.
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $0.02 |
| Output tokens | $0.02 |
Capabilities
VisionMultimodal
About PaddleOCR VL
PaddleOCR VL is a 0.9B ultra-compact vision-language model from Baidu's PaddlePaddle team for multilingual document parsing. Combines a NaViT-style dynamic resolution visual encoder with ERNIE-4.5-0.3B. Supports 109 languages for recognizing text, tables, formulas, and charts. Achieved 92.56 on OmniDocBench V1.5, surpassing larger models including DeepSeek-OCR. Released October 16, 2025.
Get Started
Model Specs
Released2025-10-16
Parameters0.9B
Context16k
ArchitectureEncoder-Decoder