ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #MoE Architecture #Model Optimization #Inference Efficiency #KV Cache

Hugging Face Releases SlimWise: Revolutionizing MoE Model Inference Efficiency and Accuracy Trade-off

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face introduces SlimWise, a serving framework that enhances the efficiency of Mixture-of-Experts (MoE) models by tailoring the expert pool to different inference phases. It performs prefill with the full model and decode with a pruned model, directly reusing the prefill-generated KV cache without conversion. Experiments on Qwen3.6-35B-A3B demonstrate that SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss. Additionally, a low-cost distilla


Background and Challenges

In modern AI systems, Mixture-of-Experts (MoE) models are popular for their ability to activate only a few experts per token. However, batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. While expert pruning can reduce this traffic, traditional approaches often also prune the compute-bound prefill phase, sacrificing model quality for little throughput benefit.

SlimWise's Innovative Solution

Hugging Face's SlimWise framework addresses these challenges by tailoring the expert pool to different inference phases. Specifically, it performs prefill with the full model and decode with a pruned model, directly reusing the prefill-generated KV cache without conversion. This approach significantly narrows the accuracy gap compared to the full model across multiple MoE backbones and pruning criteria.

Technical Highlights

  1. Training-Free KV Cache Handoff: Enables KV cache reuse between prefill and decode phases without retraining, simplifying deployment.
  2. Low-Cost Distillation Stage: Trains the decoder to continue from full-model KV caches while updating only a small subset of parameters, further optimizing generation length and accuracy.
  3. Multi-Scenario Support: Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving, making it versatile for various applications.

Experimental Results

Experiments on the Qwen3.6-35B-A3B model show that SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss. Additionally, SlimWise reveals that benchmark accuracy can conceal substantial pruning-induced changes in generation length and proposes solutions to address these distortions.

Industry Impact and Future Outlook

SlimWise provides an efficient solution for the practical deployment of MoE models, especially in resource-constrained environments. Its innovative approach not only enhances inference efficiency but also maintains model accuracy, offering new insights into AI model deployment and optimization. In the future, SlimWise is expected to be applied to more model architectures and domains, further driving the advancement of AI technology.

Developer Recommendations

  • Optimize Deployment: Leverage SlimWise's KV cache reuse mechanism to streamline the deployment of MoE models.
  • Focus on Distillation: Introduce a low-cost distillation stage during model training to optimize generation length and accuracy.
  • Explore More Applications: Experiment with applying SlimWise to other model types and tasks to explore its broader application potential.

Source: Hugging Face Daily Papers (2026-09-28)

— END —

Tags: #Hugging Face #MoE Architecture #Model Optimization #Inference Efficiency #KV Cache

Community Comments

Loading live comments and annotations…