ZICQ
中 Log in / Sign up
ZICQ Info LLMs & Foundation Models #Qwen #MoE Architecture #Inference Optimization #Large Language Models #AI Efficiency

Qwen 3.6 35B A4B+ Inference Optimization: 10.9% Latency Reduction with Zero Training Cost

Avatar of Mr.Xu

By Mr.Xu

Published: · 4 views

中文阅读 (Chinese) English Version

Summary:Qwen researchers introduce a novel optimization for sparse MoE reasoning models that dynamically adjusts the routing strategy during inference. By expanding the expert selection budget only in the later transformer layers and applying a linear decay factor, the method enhances model performance without requiring retraining. Experimental results demonstrate an 8.5% reduction in mean reasoning tokens and a 10.9% decrease in latency, while maintaining the original accuracy (84.5% vs 84.0%, statisti


A New Approach to Optimizing Sparse MoE Inference Models

The Qwen team has introduced an innovative optimization method aimed at enhancing the efficiency of sparse MoE (Mixture of Experts) inference models. The core idea is to dynamically adjust the routing strategy during the inference phase without relying on model retraining. Specifically, the method expands the expert selection budget (N≥K) only in the later transformer layers and applies a linear decay factor to the extra experts, while keeping the early layers unchanged. This adjustment upgrades the Qwen 3.6 35B A3B model to Qwen 3.6 35B A4B+.

Key Technical Highlights

  • Inference-time Optimization: No retraining or fine-tuning is required; only the routing strategy is adjusted during inference.
  • Expert Selection Budget Expansion: Expansion occurs only in the later transformer layers, with early layers remaining unchanged.
  • Linear Decay Factor: A linear decay factor is applied to the extra experts to balance resource allocation.

Experimental Results

Experiments conducted on the full MMLU-Pro (714 questions) dataset show:

  • Reduction in Reasoning Tokens: The average number of reasoning tokens decreased by 8.5%.
  • Latency Reduction: Inference latency decreased by 10.9% (p=6.5×10^-6).
  • Unchanged Accuracy: The accuracy remained at 84.5%, statistically indistinguishable from the original model's 84.0% (p=0.77).

Industry Impact and Developer Recommendations

This research demonstrates the potential of inference-time optimizations to improve model efficiency without additional training costs. For developers, this means that performance improvements can be achieved by adjusting inference strategies on existing models without the need for extensive retraining. Here are some recommendations:

  • Optimize Inference Workflows: In resource-constrained application scenarios, consider adopting similar inference-time optimization strategies.
  • Explore More Optimization Methods: Combine with other optimization techniques, such as model pruning and quantization, to further enhance model performance.
  • Focus on Model Robustness: Ensure the model's robustness across different tasks and datasets during optimization.

Future Outlook

The Qwen team plans to conduct similar experiments on larger models, such as DeepSeek V4 Flash Q2.0, to verify the method's general applicability across different models and tasks. This will provide new insights into AI model inference optimization and drive the further development of AI technology in practical applications.


Source: Reddit r/LocalLLaMA (2026-09-03)

— END —

Tags: #Qwen #MoE Architecture #Inference Optimization #Large Language Models #AI Efficiency

Community Comments

Loading live comments and annotations…