ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #LLMs & Foundation Models #Pruning #Inference Optimization #SparseDecoding

Hugging Face Releases SparseDecoding: Enhancing LLM Inference Efficiency and Accuracy

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face introduces SparseDecoding, a novel decoding-aware pruning framework that significantly enhances the efficiency and accuracy of large language model (LLM) inference. By aligning pruning objectives with decoding activations through layer-wise activation calibration and optimizing N:M sparse matrix-vector kernels, SparseDecoding achieves up to 1.48x end-to-end decoding speedup on A100 GPUs while outperforming standard fixed-text calibration on long-form generation benchmarks. This inno


Key Breakthroughs

Hugging Face's research team has introduced SparseDecoding, an innovative framework designed to address the latency issues in the decoding stage of large language model (LLM) inference. The core breakthroughs include:

  1. Decoding-Aware Pruning: By constructing calibration matrices aligned with decoding activations at the algorithmic level, SparseDecoding ensures that the pruning objective matches the activation distribution during actual decoding, addressing the distribution shift problem inherent in traditional methods.

  2. System-Level Optimization: The framework includes an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal, significantly enhancing the efficiency of sparse matrix-vector (SpMV) operations, which dominate the decoding process.

  3. Performance Improvement: Empirical results on representative LLMs such as Llama-3.1-8B, Llama-3.3-70B, and Qwen3-14B/32B demonstrate that SparseDecoding outperforms standard fixed-text calibration on long-form generation benchmarks while achieving up to 1.48x end-to-end decoding speedup on A100 GPUs.

Technical Highlights

  • Aligned Activation Distribution: Calibration matrices are built from layer-wise activations collected during dense-model autoregressive generation, ensuring consistency in activation distribution after pruning.
  • Efficient Sparse Kernel: The optimized N:M sparse matrix-vector kernel, combined with bitmask indexing and fixed-step traversal, significantly boosts the efficiency of SpMV operations.
  • Wide Applicability: The method is applicable to various LLM architectures and demonstrates excellent performance across different model scales.

Industry Impact

The release of SparseDecoding marks a significant advancement in LLM inference efficiency optimization, particularly in resource-constrained environments such as edge computing and mobile devices. The technology will help reduce the deployment costs of LLMs and facilitate their application in more domains. Additionally, the open-source nature of SparseDecoding provides researchers and developers with new tools to further advance LLM technology.

Recommendations for Developers

  • Experiment with SparseDecoding: Developers looking to deploy LLMs in resource-constrained environments should consider experimenting with SparseDecoding to improve inference efficiency.
  • Stay Updated: Hugging Face may release further optimizations and extensions for SparseDecoding, so developers should stay tuned for updates.
  • Engage with the Open-Source Community: Actively participate in the SparseDecoding open-source community, share experiences, and contribute code to collectively advance the technology.

Source: Hugging Face Daily Papers (2026-10-08)

— END —

Tags: #Hugging Face #LLMs & Foundation Models #Pruning #Inference Optimization #SparseDecoding

Community Comments

Loading live comments and annotations…