ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Neural Network Compression #LSP #Large Language Models #Model Optimization

Hugging Face Introduces LSP: Revolutionizing Neural Network Compression

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face has introduced Learnable Subspace Projections (LSP), a novel method for neural network compression. LSP learns the subspaces to discard in an end-to-end manner, significantly improving compression efficiency and model performance compared to traditional approaches. In experiments, LSP outperformed baselines across LLMs like Llama-2-7B and vision models like ViT-B/16, achieving a perplexity of 10.9 on WikiText-2 and a zero-shot accuracy of 42.2% at -70% compression for Llama-2-7B. Th


Background and Challenge

Modern Transformer models excel in natural language processing and computer vision but suffer from significant memory and computational demands, limiting their deployment in resource-constrained environments. Traditional low-rank weight factorization methods reduce memory and computation but often rely on local criteria for subspace selection, ignoring error propagation through the network, which leads to performance collapse at high compression rates.

Technical Breakthrough

Hugging Face's Learnable Subspace Projections (LSP) method addresses these challenges through the following innovations:

  • End-to-end subspace learning: LSP assigns an orthogonal projector to each linear layer or group of layers that share activations. All projectors are optimized jointly against a global objective (e.g., KL divergence or original training loss) while the pretrained weights remain frozen.
  • Initialization and rank allocation: Projectors are initialized from a whitened SVD truncation, and ranks are allocated based on the output KL induced per parameter saved.
  • Model optimization and fusion: After training, projectors merge into standard low-rank factors, with each tied group sharing one factor. In attention, this allows the model to cache one narrow latent in place of full keys and values.

Experimental Results

Experiments across multiple models demonstrate that LSP achieves a better balance between compression efficiency and performance:

  • For Llama-2-7B, at -70% compression, LSP achieves a WikiText-2 perplexity of 10.9 and a zero-shot accuracy of 42.2%, compared to 13.3 and 36.0% for the strongest baseline.
  • The factorized model decodes up to 1.6x faster than the dense model at small batch sizes.
  • At a 128k-token context, the combined memory of weights and KV cache is reduced by 13.5x, compared to at most 6.5x for traditional methods.

Industry Impact and Developer Recommendations

LSP offers a new approach to AI model compression and deployment, particularly beneficial for resource-constrained applications like mobile devices and embedded systems. Developers can apply LSP to their models to achieve more efficient compression and inference optimization. The innovation of LSP's global optimization strategy also provides new directions for future model compression research.

Future Outlook

LSP is expected to play a significant role in multimodal models, real-time AI applications, and edge computing. As AI models continue to evolve, advancements in compression technology will directly impact the adoption of AI and the expansion of its application scenarios.


Source: Hugging Face Daily Papers (2026-09-30)

— END —

Tags: #Hugging Face #Neural Network Compression #LSP #Large Language Models #Model Optimization

Community Comments

Loading live comments and annotations…