OMP-MoE: Novel Expert Pruning Framework for Efficient Mixture-of-Experts LLMs Released
By Mr.Xu
Published:
Summary:OMP-MoE is a novel training-free compression framework designed to address the memory challenges of deploying Mixture-of-Experts (MoE) based Large Language Models (LLMs). By reformulating the pruning problem as a sparse signal reconstruction task solved via Orthogonal Matching Pursuit (OMP), the method selects experts that greedily minimize reconstruction error with linear computational complexity. Experiments on models like Qwen, DeepSeek-V2, and Mixtral MoE demonstrate that OMP-MoE achieves a
Background and Challenges
Large Language Models (LLMs) have achieved remarkable success in natural language processing, but their massive parameter sizes and computational demands pose significant deployment challenges. Specifically, Mixture-of-Experts (MoE) based models, while offering scalability and efficiency advantages, suffer from high memory requirements. Existing pruning methods either incur prohibitive computational costs or neglect the dynamic interdependencies between experts, making it difficult to achieve efficient compression without compromising performance.
Core Innovations of OMP-MoE
- Orthogonal Matching Pursuit (OMP) Technique: OMP-MoE reformulates the pruning problem as a sparse signal reconstruction task and employs the OMP algorithm to select the most representative experts, minimizing reconstruction error with linear computational complexity. This approach efficiently handles large expert sets.
- Cross-Layer Expert Allocation Optimization: Utilizing a 'water-filling' strategy, OMP-MoE optimizes cross-layer expert allocation, ensuring both reconstruction quality and routing stability.
- Adaptive Inference Mechanism (OMP-MoE{\dag}): The framework introduces an adaptive inference mechanism based on energy prediction, dynamically adjusting expert activation to further enhance inference efficiency.
Experimental Results and Performance
Experiments on models like Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate that OMP-MoE achieves a 25%-50% pruning ratio while retaining 93.3% of original performance and significantly improving search and inference speeds by 33x and 1.55x, respectively. These results highlight the framework's strong capabilities in efficient compression and performance preservation.
Industry Impact and Developer Recommendations
The release of OMP-MoE provides a new solution for the efficient deployment of LLMs, particularly in resource-constrained edge devices and high-performance computing scenarios. For developers, this framework can significantly reduce model deployment costs and improve system efficiency. It is recommended that developers pay attention to the open-source code release of this technology and optimize it for their specific application scenarios.
Future Directions
As OMP-MoE gains traction, future breakthroughs are expected in the following areas:
- Multimodal Model Support: Expanding OMP-MoE's application to support pruning and compression of multimodal models.
- Automated Pruning Strategies: Developing more intelligent pruning strategies to further enhance compression efficiency and performance preservation.
- Hardware Acceleration Optimization: Combining with hardware acceleration technologies to achieve more efficient expert pruning and inference acceleration.
— END —Source: ArXiv Machine Learning (cs.LG) (2026-09-29)
Tags: #OMP-MoE #Mixture-of-Experts #Pruning Technique #LLMs & Foundation Models #Efficient Compression
Community Comments