ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Reinforcement Learning #Language Models #RP-OPD #RL

Hugging Face Introduces RP-OPD+RL Framework: A Novel Approach to Language Model Task Training

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face introduces a novel two-stage training framework called Rubric-Privileged On-Policy Distillation (RP-OPD) + RL to address the sparse training signal issue in rubric-based reinforcement learning (RL). The framework leverages rubrics as privileged teacher context for dense token-level supervision in the first stage and then uses RL to optimize rubric rewards in the second stage. Experiments on tasks like HealthBench, ResearchQA, and RubricHub Science demonstrate that this approach outp


Background and Challenge

In language model tasks, many open-ended tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this by comparing open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which specific decisions led to the final score, limiting the model's learning efficiency and performance.

Solution: RP-OPD + RL Framework

Hugging Face proposes a two-stage training framework to tackle this issue:

  1. Rubric-Privileged On-Policy Distillation (RP-OPD): In this stage, the student model, without direct access to the rubric, learns by matching the next-token distributions of a teacher model that is aware of the rubric. This method leverages the rubric as privileged context, providing denser token-level supervision.

  2. Reinforcement Learning (RL): In the second stage, RL directly optimizes the rubric rewards, allowing the model to further improve its performance beyond the distillation stage.

Experimental Results

The research team conducted experiments on tasks like HealthBench, ResearchQA, and RubricHub Science, with the following findings:

  • Performance Improvement: The RP-OPD + RL framework achieved the highest scores across all evaluated tasks, significantly outperforming methods that use only Supervised Fine-Tuning (SFT) or traditional RL.

  • Reduced Reward Hacking: Compared to the SFT + RL baseline, the RP-OPD + RL approach showed fewer signs of reward hacking on the RubricHub Science task, meaning the model did not provide irrelevant content to obtain high scores.

Technical Highlights

  • Two-Stage Training Strategy: By first performing RP-OPD and then RL, the model can more effectively utilize rubric information and improve learning efficiency.

  • Privileged Teacher Context: The RP-OPD stage uses the rubric as privileged context, providing richer learning signals to the student model.

  • Reduced Reward Hacking: This method optimizes rubric rewards while reducing strategic behaviors aimed at obtaining high scores.

Industry Impact and Developer Recommendations

  • Impact on AI Research: This framework provides a new approach to RL training based on explicit criteria, particularly suitable for tasks that require complex evaluation standards, such as healthcare and scientific research.

  • Recommendations for Developers: Developers can apply the RP-OPD + RL framework to their projects, especially when dealing with open-ended tasks and complex evaluation criteria, to enhance model performance.

  • Future Research Directions: Future research could explore how to extend the RP-OPD + RL framework to more types of tasks and optimize the two-stage training strategy for different application scenarios.


Source: Hugging Face Daily Papers (2026-10-02)

— END —

Tags: #Hugging Face #Reinforcement Learning #Language Models #RP-OPD #RL

Community Comments

Loading live comments and annotations…