HED-UCS Releases LOOM Framework: Enhancing Efficiency of Recurrent Depth Scaling in Large Language Models
By Mr.Xu
Published:
Summary:The HED-UCS team introduces LOOM, a novel framework designed to address the scaling challenges of looped Mixture-of-Experts (MoE) architectures in large language models. LOOM incorporates per-loop routers and a Looping Residual mechanism to maintain recurrent state stability while enhancing computational diversity. Experiments demonstrate that LOOM can stably scale to 9-12 loops in a 1.7B parameter model. Under near-iso-FLOP conditions, the 700M model achieves a perplexity reduction from 18.36 t
Key Breakthroughs
The HED-UCS team introduces LOOM, a framework designed to tackle the scaling challenges of looped Mixture-of-Experts (MoE) architectures in large language models. The key technical highlights of LOOM include:
- Recurrent Depth Scaling: By repeatedly applying shared Transformer blocks, LOOM increases the effective depth without increasing the parameter count.
- Stability and Diversity:
- Stability: LOOM stabilizes recurrence by scaling residual updates to limit variance growth and re-injecting the input embedding at each loop.
- Diversity: It diversifies computation through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward.
- Experimental Results:
- LOOM achieves stable scaling up to 9-12 loops in models ranging from 100M to 1.7B parameters.
- Under near-iso-FLOP conditions, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53%.
- Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%.
Industry Impact
The release of the LOOM framework offers a new technical pathway for scaling recurrent depth in large models, particularly in handling complex tasks and long-sequence modeling. Its ability to maintain computational efficiency while enhancing model performance makes it a valuable tool for AI researchers and engineers.
Developer Recommendations
- Experimental Validation: Developers are encouraged to experiment with the LOOM framework in large-scale models to validate its performance across different tasks and datasets.
- Optimization Strategies: Combine LOOM with other optimization techniques, such as mixed-precision training and model parallelism, to further improve training efficiency.
- Application Scenarios: Explore the potential of LOOM in areas such as long-sequence modeling, dialogue systems, and text generation.
Technical Highlights
- Recurrent Stability: Addresses the stability issues in recurrent depth scaling through scaling residual updates and re-injecting the input embedding.
- Computational Diversity: Introduces per-loop routers and a Looping Residual mechanism to enhance computational diversity and avoid expert selection collapse.
- Experimental Validation: Experimental results across multiple model scales validate the effectiveness and stability of the LOOM framework.
— END —Source: Hugging Face Daily Papers (2026-10-01)
Tags: #HED-UCS #LOOM #Recurrent Depth Scaling #MoE Architecture
Community Comments