ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Diffusion Models #MMD Optimization #Post-Training #DLM

Hugging Face Introduces Representation-Space MMD for Post-Training Diffusion Language Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced a novel post-training optimization method for diffusion language models (DLMs) that minimizes the Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. This approach leverages contextual features from individual token positions and optimizes the objective using policy gradients for discrete models and direct differentiation for continuous models. Experiments demonstrate improved generative perplex


Background and Challenges

Diffusion Language Models (DLMs) have shown remarkable performance in generation tasks, but their training and inference efficiency remain a key challenge. Existing optimization methods often rely on complex sampling trajectories or auxiliary models, which limit their practical applicability.

Methodology and Innovations

Hugging Face's new post-training optimization method minimizes the Maximum Mean Discrepancy (MMD) between the generated and reference distributions in the feature space of a frozen pretrained DLM. Key innovations include:

  • Feature Space Optimization: Leveraging the pretrained frozen DLM to extract contextual features and optimizing in the feature space, avoiding the need for full sampling trajectories.
  • Efficient Objective Function Optimization: Using policy gradients for discrete models and direct differentiation through generated latents for continuous models.
  • Multiple Observation Points: Retaining contextual features at individual token positions to obtain multiple observations per sequence from a single traversal pass, enhancing optimization efficiency.

Experimental Results

The experiments demonstrate the method's effectiveness in the following areas:

  • Reduced Generative Perplexity: Achieved lower generative perplexity on OpenWebText while maintaining comparable entropy.
  • Improved Accuracy-Computational Trade-off: Demonstrated better accuracy-computational trade-offs on GSM8K.
  • Enhanced Decoding Parallelism: On 16B DMax-LLaDA2.0 models, the method increased decoding parallelism while maintaining or improving accuracy on math and code benchmarks.

Industry Impact and Developer Recommendations

This method offers a novel approach to optimizing DLMs, with significant implications for:

  • Improving Model Efficiency: By avoiding complex sampling trajectories and auxiliary models, the method significantly enhances the training and inference efficiency of DLMs.
  • Boosting Model Performance: The method's strong performance across multiple benchmarks highlights its potential in complex tasks.
  • Advancing Multilingual AI: The method can be applied to multilingual DLMs, further driving the development of multilingual AI technologies.

Developers are encouraged to explore the implementation details of this method and apply it to their DLMs to improve performance.


Source: Hugging Face Daily Papers (2026-10-05)

— END —

Tags: #Hugging Face #Diffusion Models #MMD Optimization #Post-Training #DLM

Community Comments

Loading live comments and annotations…