DeepSeek
DeepSeek is the large language model family released by DeepSeek AI. In 2024-2025, V3 / R1 caused global sensation — achieving near-GPT-4 / Claude capability at extremely low training cost.
Key products
- DeepSeek V3: 671B total / 37B active (MoE), near GPT-4 on multiple benchmarks.
- DeepSeek R1: open-source reasoning model, on par with OpenAI o1.
- DeepSeek Coder / V2.5: code-specialized.
- DeepSeek-VL2: multimodal.
Technical highlights
- MLA (Multi-head Latent Attention): compresses KV cache, halving inference memory.
- MoE training framework: custom DualPipe / DeepEP for efficient parallelism.
- FP8 training: pioneer of large-scale FP8 training, further reducing cost.
- R1 distillation path: distill R1 into 1.5B-70B open-source sizes.
Influence
- Pushed global open-source LLMs to approach closed-source quality.
- Sparked debate on "is the US-China AI gap narrowing".
- MoE + FP8 + MLA training stack widely emulated by industry.