Hugging Face Releases ReSPO: Revolutionizing Sequence Policy Optimization in Reinforcement Learning
By Mr.Xu
Published:
Summary:Hugging Face has introduced ReSPO (Reshaped Sequence Policy Optimization), a novel approach to address gradient starvation in reinforcement learning. ReSPO replaces traditional clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt. This method effectively leverages long positive reasoning trajectories during early training and accelerates optimization on dense and MoE Qwen3 models, improving final trai
The Gradient Starvation Problem in Reinforcement Learning and the ReSPO Solution
In reinforcement learning, learning from verifiable rewards (RLVR) often involves reusing rollout trajectories across multiple policy updates, which increases the mismatch between the current policy and the data-generating policy. Traditional methods using clipped policy optimization suffer from a sign-dependent gradient starvation problem: clipping suppresses under-generated positive responses at the low-importance-weight tail while allowing severely over-generated negative responses to dominate the high-weight tail.
To address this, Hugging Face introduces ReSPO (Reshaped Sequence Policy Optimization). ReSPO revolutionizes sequence policy optimization through the following:
-
Two-Branch Sequence-Level Kernel: ReSPO introduces a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt.
- Positive Branch: Preserves a nonzero gradient weight for under-generated positive responses.
- Negative Branch: Suppresses heavily over-generated negative responses.
-
Advantages:
- Effectively leverages long positive reasoning trajectories during early training.
- Accelerates optimization on dense and MoE Qwen3 models.
- Improves final training scores and benchmark performance under rollout reuse.
Technical Highlights and Innovations
- Dual-Branch Kernel Design: By separating the positive and negative branches, ReSPO provides finer control over gradient updates, avoiding the limitations of traditional clipping strategies.
- α-Divergence Variational Objective: The introduction of an α-divergence variational objective allows the model to better balance exploration and exploitation during optimization.
- Exponential Variance-Control Tilt: Further stabilizes the training process and enhances model performance.
Industry Impact and Developer Recommendations
The release of ReSPO brings a new technical path to the field of reinforcement learning, particularly in handling long positive reasoning trajectories and off-policy learning. For developers, the following points are worth noting:
- Model Optimization: ReSPO can be applied to various reinforcement learning models to improve optimization efficiency and final performance.
- Resource-Constrained Scenarios: Due to its excellent performance on MoE models, ReSPO has potential applications in resource-constrained environments.
- Research Expansion: Future research can further explore the application of ReSPO in different tasks and domains, such as robotics control and autonomous driving.
Conclusion
The introduction of ReSPO marks a significant advancement in the sequence policy optimization methods of reinforcement learning. Its innovative dual-branch kernel design and α-divergence variational objective provide new ideas for solving the gradient starvation problem and have demonstrated superior performance in multiple experiments.
— END —Source: Hugging Face Daily Papers (2026-09-28)
Tags: #Hugging Face #Reinforcement Learning #ReSPO #Sequence Policy Optimization #MoE Architecture
Community Comments