tencent/Youtu-Parsing-Omni · Hugging Face
Youtu-Parsing-Omni is a compact 5B omni-modal parsing model that processes various inputs like document pages, images, charts, audio clips, or videos. It generates a structured JSON output encompassing perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The model achieves strong results at a small size, demonstrated by its state-of-the-art performance on OmniDocBench v1.6 (96.96 Overall) and competitive scores on OmniParsingBench (75.08 Avg.), making it easy to serve with included vLLM plugin and inference examples.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 9, 2026, 12:03 UTC
IngestedOffset at this time: UTC+0Oct 9, 2026, 16:00 UTC
- Published
- Oct 9, 2026, 12:03
- Ingested
- Oct 9, 2026, 16:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The output family is selected by the task prompt (--task in the examples, keys of prompts/youtu_parsing_omni.json ).
Input --task modality / subtype Key contents Document page document image / document layout elements with bbox, text / LaTeX / OTSL tables / Markdown charts / Mermaid flowcharts, reading order Natural image natural_image image / natural_image entities and text with bbox, tags, captions, global description Chart graphics_chart image / document one chart element: Markdown table, notes, caption Flowchart graphics_flowchart image / document one flowchart element: Mermaid, caption Geometry figure graphics_geometric image / document one geometric element: points, lines, arcs, shapes, geometric relations and measurements Audio audio audio / – vocal / non-vocal segments with timestamps, speakers, ASR, timbre / scene captions, acoustic events Natural video natural_video video / natural_video temporal segments with visual elements, actions, interactions, camera motion, audio track Text-rich video textrich_video video / text_rich_video segments with OCR + ASR and a Markdown structured_report of the whole video Highlights (see the technical report for details):
- Omni encoder – image, audio and interleaved audio-visual video inputs (frames + audio track) in a single model.
- Strong results at a small size – state-of-the-art on OmniDocBench v1.6 (96.96 Overall), best open-weight model on OmniParsingBench (75.08 Avg., second only to Gemini-3-Pro), and competitive with specialized models on chemical-structure (ChemOCR) and music-score (PDMX-Synth) recognition ( results ).
- Easy to serve – a vLLM plugin, pinned serving settings, task prompts and inference examples are included.