ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #LeapLabTHU #OPDVR #RLVR #Policy Distillation #Reinforcement Learning

LeapLabTHU Releases OPDVR: Combining RLVR and OPD for Optimized LLM Training

Avatar of Mr.Xu

By Mr.Xu

Published: · 14 views

中文阅读 (Chinese) English Version

Summary:The LeapLabTHU team has introduced OPDVR (On-policy Distillation with Verifiable Reward), an innovative method that seamlessly integrates Reinforcement Learning with Verifiable Rewards (RLVR) and On-policy Distillation (OPD) to address the limitations of sparse task-level feedback in RLVR and the lack of trajectory correctness in OPD. By reformulating the implicit reward and applying a ReLU gating mechanism, OPDVR ensures that correct trajectories receive non-negative rewards while preserving th


Innovation: Combining RLVR and OPD

In the training of large language models (LLMs), Reinforcement Learning (RL) and On-policy Distillation (OPD) are two widely used paradigms. RLVR provides correctness guidance through task-level feedback but suffers from sparse feedback, while OPD offers dense token-level guidance but ignores trajectory correctness, limiting the model's performance.

The OPDVR method proposed by LeapLabTHU addresses these issues through the following innovations:

  • Redefining implicit rewards: Based on trajectory correctness, the implicit reward of OPD is reformulated.
  • Introducing ReLU gating mechanism: Ensures that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, aligning the distillation signal with task success.
  • Seamless integration of RLVR and OPD: Achieves seamless integration without adding any hyperparameters.

Experimental Results: Significant Performance Improvement

OPDVR demonstrated superior performance in six reasoning benchmarks, outperforming the standard OPD method. The results indicate that OPDVR not only enhances the model's reasoning capabilities but also improves its performance in complex tasks.

Technical Highlights

  • Innovative combination of RLVR and OPD: Achieved through redefining the reward mechanism and introducing a gating mechanism.
  • No additional hyperparameters: Simplifies the training process and reduces the complexity of hyperparameter tuning.
  • Open-source code: The code is available on GitHub, allowing developers to replicate and further research.

Industry Impact and Developer Recommendations

The release of OPDVR provides new insights and methods for LLM training, particularly in applications requiring high precision and efficient inference. Developers are advised to:

  • Experiment with OPDVR: Introduce OPDVR in LLM training to enhance model performance.
  • Engage with the open-source community: Actively participate in the OPDVR open-source community to share experiences and suggestions for improvement.
  • Explore more applications: Apply OPDVR to other fields that require reinforcement learning and policy distillation, such as robotics control and autonomous driving.

Source: Hugging Face Daily Papers (2026-08-25)

— END —

Tags: #LeapLabTHU #OPDVR #RLVR #Policy Distillation #Reinforcement Learning

Community Comments

Loading live comments and annotations…