Skip to main content
ZICQ

Wiki Models & Products

DeepSeek

Models & Products
Aliases: DeepSeek V3 DeepSeek R1 ·2026-09-14

DeepSeek

DeepSeek is the large language model family released by DeepSeek AI. In 2024-2025, V3 / R1 caused global sensation — achieving near-GPT-4 / Claude capability at extremely low training cost.

Key products

  • DeepSeek V3: 671B total / 37B active (MoE), near GPT-4 on multiple benchmarks.
  • DeepSeek R1: open-source reasoning model, on par with OpenAI o1.
  • DeepSeek Coder / V2.5: code-specialized.
  • DeepSeek-VL2: multimodal.

Technical highlights

  • MLA (Multi-head Latent Attention): compresses KV cache, halving inference memory.
  • MoE training framework: custom DualPipe / DeepEP for efficient parallelism.
  • FP8 training: pioneer of large-scale FP8 training, further reducing cost.
  • R1 distillation path: distill R1 into 1.5B-70B open-source sizes.

Influence

  • Pushed global open-source LLMs to approach closed-source quality.
  • Sparked debate on "is the US-China AI gap narrowing".
  • MoE + FP8 + MLA training stack widely emulated by industry.