Xpeng AI Releases X-AuT: A Novel Framework for Efficient Compression of Speech LLM Audio Encoders
By Mr.Xu
Published: · 2 views
Summary:Xpeng AI has introduced X-AuT, a novel framework for efficiently compressing the audio encoders of speech large language models (LLMs). X-AuT selects layer combinations through short behavioral probes and restores the pruned model using representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA fine-tuning. This approach reduces inference costs while maintaining high accuracy. In tests across ten public Chinese-English benchmarks, X-AuT compressed the audi
Background and Challenges
In the application of speech large language models (LLMs), the depth of the audio encoder significantly impacts inference costs. However, removing complete encoder layers can lead to embedding perturbations, resulting in decoder errors such as deletions and premature sequence terminations.
Introduction to X-AuT Framework
Xpeng AI's X-AuT framework addresses these issues through the following techniques:
- Short Behavioral Probes: Selects optimal layer combinations to reduce model complexity.
- Representation Alignment: Ensures the pruned model aligns with the original model's representations.
- Cross-Scale Distillation: Transfers knowledge across different scales to restore model performance.
- Scheduled Student-Policy Supervision: Introduces supervision during training to optimize the student model.
- LoRA Fine-Tuning: Utilizes Low-Rank Adaptation (LoRA) for fine-tuning, further enhancing model performance.
Experimental Results
In tests across ten public Chinese-English benchmarks, X-AuT compressed the audio encoder of Qwen3-ASR-0.6B from 18 to 16 layers, reducing the macro-average error from 5.61% to 5.27%. Additionally, the 14-layer model, with a 20.7% reduction in audio tower parameters, achieved an error rate of 5.75%. Under the progressive pruning recipe, X-AuT outperformed direct pruning (5.75% vs. 6.73%).
Industry Impact
The X-AuT framework offers a new approach to optimizing speech LLMs, particularly in terms of reducing inference costs while maintaining model performance. Its applications span speech recognition, speech synthesis, and conversational systems. Developers can leverage the X-AuT framework to optimize existing speech LLM models, enhancing their performance in resource-constrained environments.
Developer Recommendations
- Model Optimization: Use the X-AuT framework to optimize existing speech LLMs to reduce inference costs.
- Experimental Validation: Validate the effectiveness of X-AuT in different application scenarios and adjust pruning strategies according to specific needs.
- Stay Updated: Keep an eye on Xpeng AI's further research advancements to gain more insights into X-AuT optimization techniques and application cases.
— END —Source: Hugging Face Daily Papers (2026-09-10)
Community Comments