Hugging Face Introduces Representation-Space MMD for Post-Training Diffusion Language Models
By Mr.Xu
Published:
Summary:Hugging Face has introduced a novel post-training optimization method for diffusion language models (DLMs) that minimizes the Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. This approach leverages contextual features from individual token positions and optimizes the objective using policy gradients for discrete models and direct differentiation for continuous models. Experiments demonstrate improved generative perplex
Background and Challenges
Diffusion Language Models (DLMs) have shown remarkable performance in generation tasks, but their training and inference efficiency remain a key challenge. Existing optimization methods often rely on complex sampling trajectories or auxiliary models, which limit their practical applicability.
Methodology and Innovations
Hugging Face's new post-training optimization method minimizes the Maximum Mean Discrepancy (MMD) between the generated and reference distributions in the feature space of a frozen pretrained DLM. Key innovations include:
- Feature Space Optimization: Leveraging the pretrained frozen DLM to extract contextual features and optimizing in the feature space, avoiding the need for full sampling trajectories.
- Efficient Objective Function Optimization: Using policy gradients for discrete models and direct differentiation through generated latents for continuous models.
- Multiple Observation Points: Retaining contextual features at individual token positions to obtain multiple observations per sequence from a single traversal pass, enhancing optimization efficiency.
Experimental Results
The experiments demonstrate the method's effectiveness in the following areas:
- Reduced Generative Perplexity: Achieved lower generative perplexity on OpenWebText while maintaining comparable entropy.
- Improved Accuracy-Computational Trade-off: Demonstrated better accuracy-computational trade-offs on GSM8K.
- Enhanced Decoding Parallelism: On 16B DMax-LLaDA2.0 models, the method increased decoding parallelism while maintaining or improving accuracy on math and code benchmarks.
Industry Impact and Developer Recommendations
This method offers a novel approach to optimizing DLMs, with significant implications for:
- Improving Model Efficiency: By avoiding complex sampling trajectories and auxiliary models, the method significantly enhances the training and inference efficiency of DLMs.
- Boosting Model Performance: The method's strong performance across multiple benchmarks highlights its potential in complex tasks.
- Advancing Multilingual AI: The method can be applied to multilingual DLMs, further driving the development of multilingual AI technologies.
Developers are encouraged to explore the implementation details of this method and apply it to their DLMs to improve performance.
— END —Source: Hugging Face Daily Papers (2026-10-05)
Tags: #Hugging Face #Diffusion Models #MMD Optimization #Post-Training #DLM
Community Comments