Youtu-Parsing-Omni

Released
2026-10-09
Last refreshed
2026-10-09
Status
Researched today
Open weightsCommercial use: conditionalMultimodalVisionJSON / Tool useMultimodalOpen Source

Youtu-Parsing-Omni is a released vision and json / tool use model with open-weight; evaluate it while provider pricing coverage matures.

Tencent AI Lab releases · 5 in the last 12 months · this family litChangelog →
Specifications
Released
2026-10-09
Parameters
5B
Architecture
Decoder Only
Specialization
multimodal
Openness
Open weights
License
Youtu-Parsing LicenseCommercial use: conditional
Weights
Available
Code
Unknown
Training
Pretrained
Created by

AI innovations for societal improvement

Shenzhen, China
Founded 2016
Website
Pricing

No tracked provider token pricing is available yet.

About

Youtu-Parsing-Omni is a 5B open-weight omni-modal parsing model from Tencent Youtu Lab, released in October 2026. One checkpoint takes a document page, natural image, chart, flowchart, geometry figure, audio clip or video with its soundtrack and returns a single structured JSON record: layout elements with bounding boxes, text, LaTeX formulas, tables, Mermaid flowcharts, speech transcripts with timestamps and speakers, sound events, camera motion and captions or reports. A single shared encoder handles images, audio and video, feeding a Youtu-LLM-style decoder. Tencent reports 96.96 overall on OmniDocBench v1.6 and 75.08 on OmniParsingBench, the best among open-weight models there.

Provider price ladder

No tracked provider token pricing is available for this model yet.

Capabilities

VisionMultimodalStructured OutputsAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

Youtu-Parsing-Omnitencent/Youtu-Parsing-Omniyoutu-parsing-omni