ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Reinforcement Learning #Length Scaling Tax #LSD #Post-Training Optimization

Hugging Face Proposes Length Self-Distillation (LSD): Optimizing Length Scaling in Reinforcement Learning Post-Training

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face's research team introduces Length Self-Distillation (LSD), a novel method to mitigate the 'length-scaling tax' (LST) in reinforcement learning (RL) post-training. LST refers to the unnecessary verbosity in responses to already-solved queries without a corresponding accuracy gain. LSD routes solved prompts to on-policy distillation while retaining the original RL objective for unsolved prompts, using an exponential moving average of the online policy as its teacher, eliminating the n


Background and Problem

The length-scaling issue in reinforcement learning (RL) post-training, known as the 'length-scaling tax' (LST), refers to the unnecessary verbosity in responses to already-solved queries. While such verbosity might indicate improved reasoning on complex problems, it is redundant and inefficient for queries that have been successfully resolved.

Method and Innovation

To address the LST problem, Hugging Face's research team proposes the Length Self-Distillation (LSD) method. The core idea of LSD is to route solved prompts to on-policy distillation while retaining the original RL objective for unsolved prompts. Specifically, LSD uses an exponential moving average of the online policy as its teacher model, eliminating the need for external models. Its main features include:

  • Policy Distillation: The outputs of solved prompts are compared with the outputs of the teacher model to optimize the model's behavior.
  • Exponential Moving Average of Online Policy: Serves as the teacher model, ensuring the distillation process is dynamic and adaptive.
  • Retention of Original RL Objective: For unsolved prompts, the original RL objective is maintained for training.

Experiments and Results

The experimental results show that LSD achieves comparable or better performance than RL across multiple variants while significantly reducing the response length growth on easy queries. Specific data include:

  • In single-turn reasoning tasks, LST decreased from 19.0% to -3.7%.
  • In multi-turn agentic tasks, LST decreased from 31.4% to 13.7%.

Industry Impact and Developer Recommendations

The LSD method provides an effective solution to the length-scaling issue in RL post-training and has the following potential application values:

  • Improving Model Efficiency: Reducing unnecessary verbosity in responses enhances the efficiency of models on simple queries.
  • Optimizing Resource Utilization: Lowering the consumption of computational and storage resources improves overall system performance.
  • Enhancing User Experience: Generating more concise and accurate responses enhances user satisfaction.

For developers, it is recommended to integrate the LSD method into existing RL training workflows to optimize the post-training performance of models. Additionally, further validation in different application scenarios is advised to assess its general applicability.


Source: Hugging Face Daily Papers (2026-09-30)

— END —

Tags: #Hugging Face #Reinforcement Learning #Length Scaling Tax #LSD #Post-Training Optimization

Community Comments

Loading live comments and annotations…