Youtu-Parsing-Omni
Youtu-Parsing-Omni is a released vision and json / tool use model with open-weight; evaluate it while provider pricing coverage matures.
- Family
- Youtu-Parsing
- Released
- 2026-10-09
- Parameters
- 5B
- Architecture
- Decoder Only
- Specialization
- multimodal
- Openness
- Open weights
- License
- Youtu-Parsing LicenseCommercial use: conditional
- Weights
- Available
- Code
- Unknown
- Training
- Pretrained
No tracked provider token pricing is available yet.
About
Youtu-Parsing-Omni is a 5B open-weight omni-modal parsing model from Tencent Youtu Lab, released in October 2026. One checkpoint takes a document page, natural image, chart, flowchart, geometry figure, audio clip or video with its soundtrack and returns a single structured JSON record: layout elements with bounding boxes, text, LaTeX formulas, tables, Mermaid flowcharts, speech transcripts with timestamps and speakers, sound events, camera motion and captions or reports. A single shared encoder handles images, audio and video, feeding a Youtu-LLM-style decoder. Tencent reports 96.96 overall on OmniDocBench v1.6 and 75.08 on OmniParsingBench, the best among open-weight models there.
Provider price ladder
No tracked provider token pricing is available for this model yet.
Capabilities
Benchmark peer barsfor Vision
No task-mapped benchmark peers are available for this model yet.
Migration checks
No linked migration route is available for this model yet.
API versions
Youtu-Parsing-Omnitencent/Youtu-Parsing-Omniyoutu-parsing-omniNo tracked provider token pricing is available yet.