ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Salt++ #Multimodal Generation #Causal Modeling #Few-Step Distillation #Xingtong Ge

Xingtong Ge Releases Salt++: A New Framework for Enhancing Few-Step Streaming Multimodal Generation

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Xingtong Ge's team introduces Salt++, a novel two-stage post-training framework designed to address the challenges of causal modeling and step distillation in few-step streaming audio-video generation. Salt++ employs Causal Self-Flow (CSF) to enhance cross-modal alignment by varying history while keeping the noisy target fixed, and Context-Aligned Autoregressive Distribution Matching Distillation (AR DMD) to match generated and reference distributions under a block-conditional KL objective. On t


Background and Challenges

Few-step streaming audio-video generation requires both causal modeling and step distillation. However, traditional training methods face two major challenges:

  1. Teacher Forcing: Pairs clean history with a noisy target but only indirectly supervises predictive contextual representations through velocity prediction.
  2. Causal Distribution Matching Distillation (DMD): Directly reusing bidirectional score models in causal generators creates a mismatch between generation and scoring contexts.

The Salt++ Framework

To address these challenges, Xingtong Ge's team introduces Salt++, a two-stage post-training framework consisting of the following two core components:

1. Causal Self-Flow (CSF)

CSF enhances cross-modal alignment by varying the history while keeping the noisy target fixed. A student model with noise-mixed history aligns its intermediate representations with those of a clean-history exponential-moving-average teacher model. This self-supervised signal encourages the student model to extract semantic information, thereby improving cross-modal alignment.

2. Context-Aligned Autoregressive DMD

Context-aligned AR DMD shares the causal mask and prefix across generator sampling, pseudo-score training, and real-score evaluation, matching generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation.

Experimental Results

On the JavisBench benchmark, Salt++ improves visual and motion quality by 57% and 45% respectively at 480p resolution compared to OmniForcing. In the extended 1664x960 generation, Salt++ outperforms bidirectional LTX-2 on six out of seven metrics.

Additionally, Salt++ introduces a scale-wise post-training stage, enabling 4-step 1664x960 generation and outperforming bidirectional LTX-2 on six reported metrics.

Industry Impact and Developer Recommendations

  • Breakthrough in Multimodal Generation: Salt++ provides an efficient training framework for few-step streaming multimodal generation, significantly enhancing generation quality and efficiency.
  • Improved Cross-Modal Alignment: The combination of CSF and AR DMD offers new insights into cross-modal alignment, aiding in the development of more powerful multimodal models.
  • Developer Recommendations: Developers are advised to pay attention to the framework design of Salt++, especially when dealing with complex multimodal tasks, and consider adopting its methods for causal modeling and step distillation.

Conclusion

The Salt++ framework, with its innovative two-stage post-training approach, addresses the key challenges in few-step streaming multimodal generation and demonstrates great potential in improving generation quality and efficiency.


Source: Hugging Face Daily Papers (2026-09-29)

— END —

Tags: #Salt++ #Multimodal Generation #Causal Modeling #Few-Step Distillation #Xingtong Ge

Community Comments

Loading live comments and annotations…