ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Xpeng AI #X-AuT #Speech LLM #Model Compression #LoRA

Xpeng AI Releases X-AuT: A Novel Framework for Efficient Compression of Speech LLM Audio Encoders

Avatar of Mr.Xu

By Mr.Xu

Published: · 2 views

中文阅读 (Chinese) English Version

Summary:Xpeng AI has introduced X-AuT, a novel framework for efficiently compressing the audio encoders of speech large language models (LLMs). X-AuT selects layer combinations through short behavioral probes and restores the pruned model using representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA fine-tuning. This approach reduces inference costs while maintaining high accuracy. In tests across ten public Chinese-English benchmarks, X-AuT compressed the audi


Background and Challenges

In the application of speech large language models (LLMs), the depth of the audio encoder significantly impacts inference costs. However, removing complete encoder layers can lead to embedding perturbations, resulting in decoder errors such as deletions and premature sequence terminations.

Introduction to X-AuT Framework

Xpeng AI's X-AuT framework addresses these issues through the following techniques:

  • Short Behavioral Probes: Selects optimal layer combinations to reduce model complexity.
  • Representation Alignment: Ensures the pruned model aligns with the original model's representations.
  • Cross-Scale Distillation: Transfers knowledge across different scales to restore model performance.
  • Scheduled Student-Policy Supervision: Introduces supervision during training to optimize the student model.
  • LoRA Fine-Tuning: Utilizes Low-Rank Adaptation (LoRA) for fine-tuning, further enhancing model performance.

Experimental Results

In tests across ten public Chinese-English benchmarks, X-AuT compressed the audio encoder of Qwen3-ASR-0.6B from 18 to 16 layers, reducing the macro-average error from 5.61% to 5.27%. Additionally, the 14-layer model, with a 20.7% reduction in audio tower parameters, achieved an error rate of 5.75%. Under the progressive pruning recipe, X-AuT outperformed direct pruning (5.75% vs. 6.73%).

Industry Impact

The X-AuT framework offers a new approach to optimizing speech LLMs, particularly in terms of reducing inference costs while maintaining model performance. Its applications span speech recognition, speech synthesis, and conversational systems. Developers can leverage the X-AuT framework to optimize existing speech LLM models, enhancing their performance in resource-constrained environments.

Developer Recommendations

  • Model Optimization: Use the X-AuT framework to optimize existing speech LLMs to reduce inference costs.
  • Experimental Validation: Validate the effectiveness of X-AuT in different application scenarios and adjust pruning strategies according to specific needs.
  • Stay Updated: Keep an eye on Xpeng AI's further research advancements to gain more insights into X-AuT optimization techniques and application cases.

Source: Hugging Face Daily Papers (2026-09-10)

— END —

Tags: #Xpeng AI #X-AuT #Speech LLM #Model Compression #LoRA

Community Comments

Loading live comments and annotations…