ZICQ
中 Log in / Sign up
ZICQ Info LLMs & Foundation Models #Tencent #Multimodal Model #Perception and Cognition #vLLM #OmniDocBench

Tencent Releases Youtu-Parsing-Omni: A Compact Omni-Modal Parsing Model for Perception and Cognition Tasks

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Tencent has released Youtu-Parsing-Omni, a compact omni-modal parsing model capable of handling diverse input formats such as document pages, natural images, charts, geometric figures, audio, and audiovisual videos. The model generates a unified structured JSON output covering both perception (layout elements, text, tables, etc.) and cognition (captions, narratives, reports, etc.) tasks. It employs a task-driven output selection mechanism and demonstrates state-of-the-art performance on benchmar


Tencent Releases Youtu-Parsing-Omni: A Compact Omni-Modal Parsing Model for Perception and Cognition Tasks

Tencent has recently released Youtu-Parsing-Omni, a compact omni-modal parsing model with the following key features:

  • Multimodal Input Handling: Supports various input formats including document pages, natural images, charts, geometric figures, audio, and audiovisual videos.
  • Unified Output Structure: Generates a unified structured JSON output covering both perception (layout elements, text, tables, etc.) and cognition (captions, narratives, reports, etc.) tasks.
  • Task-Driven Output Selection: Utilizes task prompts (--task parameter) to select the output family, such as document parsing, natural image parsing, chart parsing, etc.

Technical Highlights

  1. Unified Architecture: Employs a unified architecture to handle seven parsing tasks, simplifying the model structure and enhancing processing efficiency.
  2. Omni Encoder: Integrates image, audio, and audiovisual frame encoding capabilities, enabling multimodal input processing within a single model.
  3. Outstanding Performance: Achieves a 96.96 overall score on OmniDocBench v1.6, making it the leading open-weight model. On OmniParsingBench, it scores an average of 75.08, second only to Gemini-3-Pro.
  4. Ease of Deployment: Provides a vLLM plugin, preset serving settings, task prompts, and inference examples, simplifying the practical application of the model.

Industry Impact

The release of Youtu-Parsing-Omni marks a significant advancement in the field of multimodal AI models for handling complex tasks. Its powerful perception and cognition capabilities make it highly applicable in areas such as document processing, image recognition, and audio analysis. For instance, in scenarios like enterprise document parsing, automated report generation, and cross-modal data fusion, Youtu-Parsing-Omni can significantly boost efficiency. Additionally, its ease of deployment lowers the barrier to entry for developers, promoting the widespread adoption of multimodal AI technology.

Developer Recommendations

  • Explore Multimodal Application Scenarios: Developers can leverage Youtu-Parsing-Omni to process diverse input formats and explore new application scenarios, such as intelligent document processing and cross-modal data analysis.
  • Optimize Task Prompts: By adjusting the task prompt parameters, developers can customize the model's output content to meet specific requirements.
  • Combine with Other Tools: Integrating Youtu-Parsing-Omni with other AI tools, such as natural language processing models or knowledge graphs, can further enhance application effectiveness.

Source: Reddit r/LocalLLaMA (2026-10-09)

— END —

Tags: #Tencent #Multimodal Model #Perception and Cognition #vLLM #OmniDocBench

Community Comments

Loading live comments and annotations…