Hugging Face Proposes New Method for Optimizing Looped Models: Significant Improvements in Training and Inference Effici
By Mr.Xu
Published:
Summary:Hugging Face's research team has introduced a novel method for optimizing looped language models, enhancing efficiency in training, decoding, prefill, and reinforcement learning (RL). The approach enables truncated backpropagation, terminal key-value (KV) sharing with minimal accuracy loss, and 1.79x faster prefill for distilled students. Additionally, RL updates computed from saved rollout states are 2x faster than backpropagating through the replayed trajectory. By improving depth priors and o
A New Breakthrough in Optimizing Looped Language Models
Hugging Face's research team has introduced a novel optimization method for looped language models, addressing efficiency bottlenecks in training, decoding, prefill, and reinforcement learning (RL). The key innovations of this approach include:
1. Fixed Point Inference and Truncated Backpropagation
As recurrent states approach fixed points, the influence of the path diminishes. This property enables the use of truncated backpropagation during training, reducing computational costs.
2. Terminal Key-Value (KV) Sharing
The terminal KV sharing mechanism allows the model to utilize KV caches efficiently during decoding with almost no loss in accuracy. This not only speeds up inference but also reduces memory usage.
3. Prefill Acceleration for Distilled Students
The method enables a 1.79x faster prefill for distilled student models while maintaining high-precision outputs. This is particularly important for applications requiring rapid responses.
4. RL Update Acceleration
By computing gradients from saved rollout states, RL updates are 2x faster than traditional backpropagation methods. This significantly improves the efficiency of reinforcement learning training.
5. Improvements in Depth Prior and Orthogonal Injection
Existing methods have limitations in depth priors and input injection. Hugging Face's new approach learns priors from prediction feedback and introduces orthogonal injection to eliminate interference from the input component, thereby reducing perplexity across multiple scales.
Technical Highlights
- KV Cache Efficiency: The KV cache size is reduced by 3x while maintaining performance comparable to the full cache.
- Training and Inference Speed: The prefill speed for distilled students is increased by 1.79x, and RL update speed is doubled.
- Perplexity Reduction: The new method reduces perplexity at every scale from 100M to 1.6B parameters.
Industry Impact
This research provides new insights into the optimization of looped language models, particularly in resource-constrained and real-time application scenarios. For example, in autonomous driving, real-time translation, and intelligent assistants, this method can significantly enhance the practical application efficiency of the model.
Developer Recommendations
- Focus on KV Cache Management: Developers can leverage the terminal KV sharing mechanism to optimize memory usage of the model.
- Explore Orthogonal Injection: Introducing orthogonal injection in model training can improve training efficiency and model performance.
- Combine with Reinforcement Learning: Utilizing the RL update acceleration method can train reinforcement learning models more efficiently.
Conclusion
Hugging Face's new method offers an innovative solution for optimizing looped language models, demonstrating significant potential in training and inference efficiency. This research not only advances AI technology but also opens up new possibilities for AI application scenarios.
— END —Source: Hugging Face Daily Papers (2026-10-05)
Tags: #Hugging Face #Looped Models #Optimization #Training Efficiency #Inference Efficiency
Community Comments