ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Diffusion Models #Image Generation #AI Image Generation #CompVis

CompVis Introduces Improved Distributional Diffusion Models: Breakthrough in Image Generation Efficiency and Quality

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:CompVis introduces an improved Distributional Diffusion Model (iDDM) that addresses key limitations in scaling diffusion models for modern image generation. By deferring particle expansion to later transformer layers and introducing time-dependent scoring rule schedules, the model achieves state-of-the-art performance on ImageNet-256² with 4.48 FID at 4 steps and 2.38 FID at 50 steps using DiT-XL/2. This advancement eliminates the need for teacher models, self-distillation, or Jacobian-vector pr


Background and Challenges

Diffusion Models have made significant strides in image generation, but scaling them to modern image generation tasks faces two key challenges:

  1. High overhead of multi-particle training: As the number of particles increases, the training overhead grows linearly, limiting the model's scalability.
  2. Globally fixed scoring rule hyperparameters: Diffusion models use globally fixed scoring rule hyperparameters, resulting in a single trade-off across sampling budgets, which is inflexible for different scenarios.

Innovations

CompVis introduces the Improved Distributional Diffusion Model (iDDM) to address these challenges through the following methods:

  1. Deferred particle expansion: Delaying the particle expansion operation to later Transformer layers reduces computational overhead during training.
  2. Time-dependent scoring rule schedules: Introducing time-dependent scoring rule schedules based on the dynamical mechanisms of Biroli2024 replaces globally fixed hyperparameters, enabling more flexible trade-off optimization.

Combined with a DiT-based latent space setup, iDDM achieves the following performance on ImageNet-256²:

  • 4-step generation: FID of 4.48.
  • 50-step generation: FID of 2.38.

This result comes from a single model trained from scratch, without the need for teacher models, self-distillation, or Jacobian-vector products (JVP), demonstrating its efficiency and high-quality output in image generation tasks. Additionally, the approach successfully transfers to text-to-image generation, showcasing its broad application potential.

Key Features

  • Deferred particle expansion: Reduces training overhead and enhances scalability.
  • Time-dependent scoring rule schedules: Enables more flexible trade-off optimization.
  • No additional components needed: Eliminates the need for teacher models, self-distillation, or JVP, simplifying the model.
  • Cross-task transferability: Successfully applied to text-to-image generation, highlighting its versatility.

Industry Impact and Developer Recommendations

iDDM represents a significant advancement in AI image generation, particularly in handling complex image generation tasks. Its efficient training process and high-quality output provide developers with a more powerful tool. Developers should consider the following:

  • Model optimization: Utilize deferred particle expansion and time-dependent scoring rule schedules to optimize the training process of existing diffusion models.
  • Cross-task application: Explore the potential of iDDM in text-to-image generation and other multimodal tasks.
  • Resource management: The low-overhead nature of iDDM makes it ideal for resource-constrained environments.

Conclusion

CompVis's iDDM model marks an important breakthrough in the image generation field, providing a new direction for the further development of AI technology. Its efficient training process and high-quality output demonstrate the great potential of diffusion models in modern image generation tasks.


Source: Hugging Face Daily Papers (2026-09-29)

— END —

Tags: #Diffusion Models #Image Generation #AI Image Generation #CompVis

Community Comments

Loading live comments and annotations…