arXiv Releases SharedSAE: A New Method for Sharing Feature Dictionaries Across Language Models
By Mr.Xu
Published: · 6 views
Summary:arXiv introduces SharedSAE, a novel approach that replaces per-model sparse autoencoders (SAEs) with a single shared SAE across language models. By combining a shared dictionary with model-specific encoder-decoder pairs, SharedSAE achieves 96.6% of the mean explained variance of dedicated SAEs while exhibiting 1.8 times higher cross-model correlations in latent activations. This method enhances training efficiency and allows new models to adapt quickly to the shared latent descriptions, offering
Key Breakthroughs
- Shared Sparse Autoencoder (SharedSAE): SharedSAE introduces a shared dictionary and model-specific encoder-decoder pairs, replacing the traditional per-model SAEs.
- Performance Benefits: It retains 96.6% of the mean explained variance of dedicated SAEs while achieving 1.8 times higher cross-model correlations in latent activations.
- Model Transferability: The latent descriptions of SharedSAE can be transferred across models, significantly improving training efficiency.
- Rapid Adaptation for New Models: After freezing the dictionary, new models can efficiently adapt to the shared dictionary while maintaining near-dedicated-SAE reconstruction quality.
Technical Analysis
The core innovation of SharedSAE lies in its shared dictionary design, which captures common features across different language models while preserving each model's uniqueness through model-specific encoder-decoder pairs. Compared to traditional methods, SharedSAE avoids the repetitive training of SAEs, significantly reducing computational costs. Additionally, during inference, SharedSAE only requires model-specific dropout mechanisms without relying on all models, thus enhancing inference efficiency.
Industry Impact
- Enhanced Training Efficiency: By sharing the dictionary, SharedSAE reduces the need for repetitive SAE training, saving computational resources.
- Accelerated New Model Development: New models can quickly adapt to the shared dictionary, shortening the development cycle.
- Facilitated Cross-Model Collaboration: The cross-model transferability of SharedSAE opens new possibilities for multi-model collaboration.
Recommendations for Developers
- Experiment with SharedSAE: For projects involving multiple language models, consider using SharedSAE to improve training and inference efficiency.
- Assess Model Compatibility: When applying SharedSAE, evaluate the compatibility of new models with the shared dictionary to ensure reconstruction quality.
- Explore More Applications: The cross-model nature of SharedSAE makes it promising for multi-model collaboration and knowledge transfer scenarios, warranting further exploration.
— END —Source: ArXiv Machine Learning (cs.LG) (2026-09-07)
Tags: #SharedSAE #Large Language Models #Feature Representation #arXiv #Sparse Autoencoders
Community Comments