Daedalus-150M Released: A Convolution-Attention Hybrid Architecture for CPU Inference
By Mr.Xu
Published:
Summary:Daedalus-150M is a lightweight language model specifically optimized for CPU inference, featuring an innovative hybrid architecture that balances efficiency and performance. By retaining full attention in only 6 out of 18 blocks and using short convolutions for the rest, the model reduces memory usage and enhances inference speed. Trained on 59.9B tokens, it achieves a score of 47.31 on a five-task benchmark, outperforming models like GPT-2, Pythia, OPT, and GPT-Neo, and even surpassing the larg
Innovative Architecture Design
Daedalus-150M features a unique hybrid architecture that combines full attention mechanisms with short convolutions to cater to CPU inference needs. Specifically, the model retains full attention in only 6 out of 18 blocks, while the remaining 12 blocks utilize short convolutions. This design significantly reduces memory usage while maintaining high inference speed.
Training and Performance
The model was trained from scratch on 59.9B tokens and achieved a score of 47.31 on a five-task benchmark, outperforming models like GPT-2 124M, Pythia-160M, OPT-125M, and GPT-Neo-125M, which were trained on three to six times more data. Additionally, despite MobileLLM-125M being trained on 1 trillion tokens, Daedalus-150M demonstrates superior performance in terms of efficiency and effectiveness.
Technical Highlights
- Hybrid Architecture: By combining full attention and short convolutions, Daedalus-150M maintains efficient inference while significantly reducing memory usage.
- Efficient Inference: At a context length of 2048 tokens, Daedalus-150M decodes 1.76 times faster than comparable models.
- Lightweight Design: 4-bit weights and optimized memory management make it suitable for standard CPU environments.
Limitations and Challenges
Although Daedalus-150M excels in several aspects, the research team also acknowledges its limitations, such as higher 4-bit quality costs, inactive convolution channels, and an overly large vocabulary.
Industry Impact and Recommendations
The release of Daedalus-150M provides a new solution for CPU inference scenarios, particularly for resource-constrained devices. Its lightweight design and efficient performance make it a promising candidate for applications in edge computing, IoT, and mobile devices. Developers are encouraged to deploy the model in specific application scenarios and fine-tune it according to their needs to achieve optimal performance.
— END —Source: Hugging Face Daily Papers (2026-08-20)
Tags: #Daedalus #CPU Inference #Hybrid Architecture #Lightweight Model #Efficient Inference
Community Comments