ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #CoEvoWhen #Long-Video Processing #VLM #Policy-Tool Coevolution

Hugging Face Releases CoEvoWhen: Revolutionizing Ultra-Long Video Temporal Grounding

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced CoEvoWhen, a novel framework that enhances ultra-long video temporal grounding through the coevolution of policies and tools. By leveraging the reasoning trajectories of a Vision-Language Model (VLM), CoEvoWhen jointly optimizes high-level policies and executable media tools without requiring model parameter updates, forming reusable skills. Experiments across five benchmarks demonstrate that CoEvoWhen significantly improves temporal grounding accuracy in ultra-long v


Core Breakthroughs

The CoEvoWhen framework introduced by Hugging Face revolutionizes ultra-long video temporal grounding through the following innovations:

  1. Policy-Tool Coevolution: CoEvoWhen leverages the reasoning trajectories of a Vision-Language Model (VLM) to jointly optimize high-level policies and executable media tools, forming reusable skills without updating model parameters.
  2. External Skill Updater: An external skill updater distills transferable task experience, refining the orchestration of long-range image-based and fine-grained video-based observations, and employs its coding capabilities to upgrade or create tools for long-video evidence acquisition.
  3. Autonomous Tool Orchestration: Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model.

Experimental Results

Experiments across five benchmarks and three VLMs demonstrate that CoEvoWhen:

  • Improves Temporal Grounding Accuracy: Significantly enhances accuracy in ultra-long video temporal grounding tasks.
  • Reduces Visual Token Costs: Achieves a notable reduction in visual token costs during inference.
  • Demonstrates Generalizability: The evolved skill yields substantial performance gains on general long-video QA tasks without additional task-specific evolution.

Technical Highlights

  • No Model Parameter Updates: The framework enhances performance without the need for model parameter updates through policy-tool coevolution.
  • Reusable Skills: The skills evolved are reusable across related tasks, showcasing their generalizability.
  • Long-Video Understanding: CoEvoWhen is particularly effective for long-video understanding tasks, addressing the balance between long-range evidence search and fine-grained event understanding.

Industry Impact and Developer Recommendations

The release of CoEvoWhen provides new avenues for the long-video processing field, especially in applications requiring efficient processing and understanding of large volumes of video data, such as video surveillance, film production, and online education. Developers can integrate CoEvoWhen into existing VLM workflows to enhance the performance of long-video processing tasks. Additionally, the coevolution mechanism of CoEvoWhen offers insights for optimizing intelligent agent systems in other domains.


Source: Hugging Face Daily Papers (2026-09-30)

— END —

Tags: #Hugging Face #CoEvoWhen #Long-Video Processing #VLM #Policy-Tool Coevolution

Community Comments

Loading live comments and annotations…