ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Reinforcement Learning #Low-Precision Execution #TRIAGE #Qwen Models

Hugging Face Releases TRIAGE: Revolutionizing Low-Precision Reinforcement Learning Optimization

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced TRIAGE, a novel method aimed at addressing the instability in policy optimization caused by low-precision execution in reinforcement learning (RL). TRIAGE employs segment-level diagnosis and selective rebalancing of policy-gradient updates while maintaining native NVFP4 4-bit weight and activation forward execution. Experiments on Qwen3-4B and Qwen3-30B-A3B models demonstrate that TRIAGE achieves full-precision performance across five mathematical reasoning benchmarks


Background and Challenges

Low-precision execution is a crucial technique for accelerating reinforcement learning (RL) training for large language models. However, the discrepancies it introduces can destabilize policy optimization. Existing methods struggle to effectively identify and address the impact of such instability on the training process.

Overview of TRIAGE

Hugging Face's TRIAGE addresses these challenges through the following approaches:

  • Direction-Aware Stability Optimization: TRIAGE identifies locally amplifying and contracting update contributions in the policy gradient direction, rather than relying solely on mismatch magnitude.
  • Segment-Level Diagnosis and Rebalancing: By diagnosing at the segment level, TRIAGE can identify and rebalance the updates that negatively impact the training process, thereby stabilizing the optimization.
  • Maintaining Native NVFP4 Execution: TRIAGE modifies the optimization objective while retaining native NVFP4 4-bit weight and activation forward execution to adapt to the low-precision execution environment.

Experimental Results

Experiments on Qwen3-4B and Qwen3-30B-A3B models demonstrate that:

  • Training Stability: TRIAGE maintains a stable optimization process throughout the training horizon.
  • Performance Improvement: TRIAGE achieves full-precision performance across five mathematical reasoning benchmarks.
  • Throughput Enhancement: TRIAGE provides up to 2.3x higher rollout throughput compared to BF16.

Industry Impact and Developer Recommendations

The release of TRIAGE offers new optimization strategies for low-precision reinforcement learning, particularly in scenarios where resource efficiency is critical but high performance is required. For developers, the following points are noteworthy:

  • Optimization Strategy: TRIAGE demonstrates the potential of fine-grained diagnosis and rebalancing strategies to enhance training stability.
  • Resource Efficiency: The method significantly reduces computational resource requirements while maintaining high performance.
  • Application Scenarios: Suitable for large-scale RL model development where fast inference and efficient training are essential.

Technical Highlights

  • Direction-Aware Updates: By identifying local amplification and contraction in the policy gradient direction, TRIAGE enables more precise optimization.
  • Segment-Level Diagnosis Mechanism: Processing update contributions in segments allows for the identification and repair of potential instability factors.
  • Retention of Native Execution: Optimizing without affecting native NVFP4 execution enhances the method's practicality.

Source: Hugging Face Daily Papers (2026-10-05)

— END —

Tags: #Hugging Face #Reinforcement Learning #Low-Precision Execution #TRIAGE #Qwen Models

Community Comments

Loading live comments and annotations…