ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Low-bit Quantization #Attention Mechanism #Video Generation #AI Inference Optimization

Hugging Face Introduces VC-Attention: Breakthrough Optimization for Low-bit Attention Mechanisms

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced VC-Attention, a novel low-bit attention framework designed to address the performance bottlenecks of attention mechanisms in tasks like video generation with long sequences. By employing Value Smoothing (V-Smooth) and fused probability casting (ExpCast-FP8), VC-Attention significantly enhances the accuracy and speed of low-bit quantization. In multiple benchmarks, VC-Attention outperforms existing low-bit baselines by 1.13-1.19x in fidelity and accelerates the attenti


Technical Breakthroughs and Core Innovations

The core innovations of the VC-Attention framework include:

  • Value Smoothing (V-Smooth): This technique reorders value tokens through lightweight online clustering, allowing tokens within a hardware block to be quantized more effectively together. It quantizes only the residual after subtracting the block mean and restores the mean from the row sum maintained by the online Softmax, thereby reducing output errors.

  • Probability Casting (ExpCast-FP8): This method maps log-domain scores directly to E4M3 probability codes using a single fused multiply-add operation, eliminating the FP32 exponential operation and format conversion. This enhances computational efficiency and reduces hardware resource consumption.

Performance

VC-Attention demonstrates superior performance in multiple benchmarks:

  • Accuracy Improvement: Across datasets like Wan2.2, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3, VC-Attention outperforms low-bit baselines by 1.13-1.19x in fidelity.

  • Inference Speed: On datacenter Blackwell and Hopper GPUs, VC-Attention accelerates the inference speed by 1.46-1.59x compared to BF16 FlashAttention-4. On workstation GPUs, the speed is increased by 2.3-3.6x.

  • End-to-End Speed: The generation of clips is 1.13-1.19x faster, and the end-to-end speed is improved by 1.36-1.70x.

Industry Impact and Developer Recommendations

The release of VC-Attention opens new possibilities for AI models in scenarios requiring efficient inference and low resource consumption, particularly in tasks like video generation and real-time processing where latency and accuracy are critical. Developers should consider the following:

  • Optimizing Low-Bit Quantization: VC-Attention presents a new path for low-bit quantization. Developers can explore its application in other scenarios that demand efficient inference.

  • Hardware Adaptation: VC-Attention has been implemented across various hardware platforms. Developers can further optimize and extend it based on their specific needs.

  • Cross-Domain Applications: While initially targeted at video generation, this technology can be applied to other domains that involve processing long sequences of data, such as natural language processing and audio processing.

Conclusion

The introduction of VC-Attention marks a significant advancement in low-bit attention mechanisms, providing a new technical pathway for AI models in scenarios requiring efficient inference and low resource consumption.


Source: Hugging Face Daily Papers (2026-09-14)

— END —

Tags: #Hugging Face #Low-bit Quantization #Attention Mechanism #Video Generation #AI Inference Optimization

Community Comments

Loading live comments and annotations…