ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #MiniMax #Video Generation #Hybrid Attention #VDN #AI Acceleration

MiniMax Releases Video DeltaNet: Revolutionizing Livestream Video Generation Efficiency

Avatar of Mr.Xu

By Mr.Xu

Published: · 4 views

中文阅读 (Chinese) English Version

Summary:MiniMax has introduced Video DeltaNet (VDN), a hybrid attention mechanism model optimized for livestream video generation. VDN combines local Softmax attention with bidirectional linear memory to handle long-range video context. Its linear branch introduces Video Delta Attention (VDA), updating memory once per frame by incorporating spatial tokens, while separate output projections and learnable gates calibrate the two branches. VDN-H3, instantiated on MiniMax H3, achieves a 14.5x speedup over t


Background and Challenges

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. While linear attention has been widely adopted in large language models, directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation.

Innovations of Video DeltaNet (VDN)

  1. Hybrid Attention Mechanism: VDN combines local Softmax attention with bidirectional linear memory to handle long-range video context.
  2. Video Delta Attention (VDA): The linear branch of VDN introduces VDA, updating memory once per frame by incorporating spatial tokens, while maintaining fine-grained interactions.
  3. Learnable Gating Mechanism: VDN uses learnable gates to calibrate the outputs of the two branches, ensuring the model's adaptability to different tasks.
  4. Staged Teacher-Alignment Recipe: VDN employs a staged teacher-alignment recipe to gradually introduce the new pathway into pretrained models, ensuring stable and efficient training.

Implementation and Performance

VDN is instantiated on MiniMax H3 and applied to video-to-video interactions, while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 achieves a 14.5x speedup over the 50-step dense H3 baseline, processing a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs.

Industry Impact and Developer Recommendations

The release of VDN marks a significant breakthrough in the field of video generation for handling long-range spatiotemporal dependencies. Its efficient performance and innovative architecture provide new solutions for video generation tasks, particularly for scenarios requiring real-time processing, such as livestream video generation. Developers can refer to VDN's architecture to optimize the performance of existing video generation models and explore its potential applications in other domains.

Future Outlook

With the introduction of VDN, future research can further explore its applications in multimodal interactions and real-time video editing. Additionally, the architectural design of VDN offers new insights for other types of diffusion models, potentially driving advancements in the broader AI video generation field.


Source: Hugging Face Daily Papers (2026-09-17)

— END —

Tags: #MiniMax #Video Generation #Hybrid Attention #VDN #AI Acceleration

Community Comments

Loading live comments and annotations…