Xingtong Ge Releases Salt++: A New Framework for Enhancing Few-Step Streaming Multimodal Generation
By Mr.Xu
Published:
Summary:Xingtong Ge's team introduces Salt++, a novel two-stage post-training framework designed to address the challenges of causal modeling and step distillation in few-step streaming audio-video generation. Salt++ employs Causal Self-Flow (CSF) to enhance cross-modal alignment by varying history while keeping the noisy target fixed, and Context-Aligned Autoregressive Distribution Matching Distillation (AR DMD) to match generated and reference distributions under a block-conditional KL objective. On t
Background and Challenges
Few-step streaming audio-video generation requires both causal modeling and step distillation. However, traditional training methods face two major challenges:
- Teacher Forcing: Pairs clean history with a noisy target but only indirectly supervises predictive contextual representations through velocity prediction.
- Causal Distribution Matching Distillation (DMD): Directly reusing bidirectional score models in causal generators creates a mismatch between generation and scoring contexts.
The Salt++ Framework
To address these challenges, Xingtong Ge's team introduces Salt++, a two-stage post-training framework consisting of the following two core components:
1. Causal Self-Flow (CSF)
CSF enhances cross-modal alignment by varying the history while keeping the noisy target fixed. A student model with noise-mixed history aligns its intermediate representations with those of a clean-history exponential-moving-average teacher model. This self-supervised signal encourages the student model to extract semantic information, thereby improving cross-modal alignment.
2. Context-Aligned Autoregressive DMD
Context-aligned AR DMD shares the causal mask and prefix across generator sampling, pseudo-score training, and real-score evaluation, matching generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation.
Experimental Results
On the JavisBench benchmark, Salt++ improves visual and motion quality by 57% and 45% respectively at 480p resolution compared to OmniForcing. In the extended 1664x960 generation, Salt++ outperforms bidirectional LTX-2 on six out of seven metrics.
Additionally, Salt++ introduces a scale-wise post-training stage, enabling 4-step 1664x960 generation and outperforming bidirectional LTX-2 on six reported metrics.
Industry Impact and Developer Recommendations
- Breakthrough in Multimodal Generation: Salt++ provides an efficient training framework for few-step streaming multimodal generation, significantly enhancing generation quality and efficiency.
- Improved Cross-Modal Alignment: The combination of CSF and AR DMD offers new insights into cross-modal alignment, aiding in the development of more powerful multimodal models.
- Developer Recommendations: Developers are advised to pay attention to the framework design of Salt++, especially when dealing with complex multimodal tasks, and consider adopting its methods for causal modeling and step distillation.
Conclusion
The Salt++ framework, with its innovative two-stage post-training approach, addresses the key challenges in few-step streaming multimodal generation and demonstrates great potential in improving generation quality and efficiency.
— END —Source: Hugging Face Daily Papers (2026-09-29)
Tags: #Salt++ #Multimodal Generation #Causal Modeling #Few-Step Distillation #Xingtong Ge
Community Comments