ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Reinforcement Learning #Flow Models #MEND

Hugging Face Releases MEND: A Novel Reinforcement Learning Algorithm Based on Proximal Velocity Matching

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced MEND, a novel reinforcement learning algorithm designed to enhance the reward post-training process for flow models. MEND caps rewards within each prompt group and only proposes moves for samples below the cap, accepting them only when the capped reward gain exceeds a quadratic displacement cost. This approach outperforms Flow-GRPO in 100 updates and surpasses ReFL and DiffusionNFT under an equal-budget protocol, demonstrating its efficiency and versatility.


Key Breakthroughs

Hugging Face has introduced MEND (Proximal Velocity Matching for Flow Models via Reinforcement Learning), a novel reinforcement learning algorithm aimed at improving the reward post-training process for flow models. The key technical highlights of MEND include:

  • Reward Capping Mechanism: MEND caps rewards within each prompt group to ensure that well-performing samples do not undergo unnecessary adjustments, thereby enhancing training efficiency.
  • Reward Gradient-Based Move Proposals: For samples below the cap, MEND proposes moves along the reward gradient and only accepts them when the capped reward gain exceeds a quadratic displacement cost.
  • No KL Penalty Term: Unlike traditional flow model training methods based on KL penalties, MEND does not require a KL term or a frozen reference model, simplifying the training process.
  • Performance Advantages: In 100 updates, MEND outperforms Flow-GRPO, which requires thousands of updates, on five out of six evaluators. Under an equal-budget protocol, MEND surpasses ReFL and DiffusionNFT across four training rewards, demonstrating its efficiency and versatility.

Technical Analysis

The core innovation of MEND lies in its proximal velocity matching approach, which limits the range and magnitude of reward adjustments to avoid the over-adjustment problems common in traditional methods. This approach not only improves training efficiency but also reduces computational resource consumption, making it more practical in resource-constrained environments.

Industry Impact

The release of MEND provides a more efficient and flexible method for the reward post-training process of flow models. Its versatility makes it applicable to any flow model backend with a differentiable reward, offering AI researchers and engineers a more powerful tool. Here are some potential impacts of MEND on the industry:

  • Enhanced Model Training Efficiency: MEND significantly improves model training efficiency by reducing unnecessary sample adjustments and simplifying the training process.
  • Reduced Computational Costs: Since MEND does not require a KL penalty term or a frozen reference model, the computational costs during training are reduced.
  • Promotion of Flow Model Applications: The versatility and efficiency of MEND will promote the application of flow models in more fields, such as image generation, audio processing, and natural language processing.

Developer Recommendations

For AI developers, MEND offers a new reinforcement learning training method. Here are some recommendations:

  • Apply MEND to Existing Flow Models: Developers can try applying MEND to existing flow model post-training processes to verify its performance improvement effects.
  • Explore MEND Applications in Other Fields: Due to its versatility, developers can explore MEND's applications in other fields, such as robotics control, recommendation systems, and autonomous driving.
  • Combine Other Technologies to Optimize MEND: Developers can combine other technologies (such as knowledge distillation, quantization, etc.) to further optimize MEND's performance.

Source: Hugging Face Daily Papers (2026-10-05)

— END —

Tags: #Hugging Face #Reinforcement Learning #Flow Models #MEND

Community Comments

Loading live comments and annotations…