ZICQ
中 Log in / Sign up
ZICQ Info Open Source AI #Sparse Memory LM #Open-Source Model #Memory Mapping #Triton Kernels #AI Inference

Open-Sourced Sparse Memory LM: 21M Parameter Model with 6.4B Parameter Lookup Table Matches 114M Dense Model

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A Reddit user released an open-source model named Sparse Memory LM, which combines a small 21M parameter model with a 6.4B parameter lookup table, achieving performance comparable to a 114M dense model. The model leverages memory mapping to store the lookup table on an NVMe SSD, requiring only 0.4GB of VRAM and achieving a token generation rate of 140 tokens per second. While the generated text is fluent, it occasionally produces factually incorrect information. The researcher also attempted to


Key Technical Highlights

  1. Model Architecture and Innovation

    • Sparse Memory Mechanism: The Sparse Memory LM introduces a 16.8M-row lookup table (6.4B parameters in total) to combine the strengths of a small model with the performance of a larger model.
    • Memory Mapping Technique: The lookup table is stored on an NVMe SSD and accessed via memory mapping, eliminating the need for VRAM and requiring only 0.4GB of VRAM to run.
  2. Performance

    • Model Performance: The 21M parameter model with the lookup table matches the performance of a 114M dense model when trained on the same 500M Wikipedia tokens.
    • Inference Speed: The model generates approximately 140 tokens per second on an RX 9070 GPU.
  3. Limitations

    • Text Generation Issues: While the generated text is fluent, it occasionally produces factually incorrect information.
    • Scalability Challenges: Applying the sparse memory mechanism to larger models like Qwen3.5-0.8B did not yield significant improvements.
  4. Technical Implementation Details

    • Triton Kernels: The researcher wrote Triton kernels to ensure seamless operation across different hardware, including Radeon, MI350X, and H100/H200.
    • Open-Source Code and Tools: The model and tools are open-sourced, with the code hosted on GitHub and an interactive visualization tool provided for users to explore the lookup table entries accessed by the model.

Industry Impact and Developer Recommendations

  • Impact on AI Research: Sparse Memory LM demonstrates the potential of combining large lookup tables with small models, offering new possibilities for AI applications in resource-constrained environments.
  • Hardware Requirements: The model has low VRAM requirements but high SSD read speed demands, requiring developers to balance hardware configurations based on specific application scenarios.
  • Future Research Directions: Further optimization of the lookup table structure and access mechanism, or exploring the combination of sparse memory mechanisms with larger models, could enhance model performance and the accuracy of generated text.

Conclusion

Sparse Memory LM is an innovative research project that showcases the potential of efficient AI inference in resource-constrained environments. Despite some limitations, its open-source nature and technical implementation details provide valuable insights and a solid foundation for the AI research community.


Source: Reddit r/LocalLLaMA (2026-10-06)

— END —

Tags: #Sparse Memory LM #Open-Source Model #Memory Mapping #Triton Kernels #AI Inference

Community Comments

Loading live comments and annotations…