Skip to main content
ZICQ

Wiki Concepts

Pre-training

Concepts
Aliases: pretrain ·2026-09-14

Pre-training

Pre-training is the first stage of LLM training: using massive unlabeled text (trillion-token scale), the model learns language knowledge and world knowledge through the self-supervised "next-token prediction" task.

Typical scale

  • GPT-2: ~10B tokens
  • LLaMA 1: ~1.4T tokens
  • LLaMA 3: ~15T tokens
  • Top industry models: tens of T tokens and up

Data sources

  • Web pages: Common Crawl is the largest single source (60-80%), needs strict cleaning.
  • Code: GitHub, StackOverflow (improve code ability).
  • Books: high-quality long text (Project Gutenberg, Books3).
  • Academic: arXiv papers, PubMed (improve scientific reasoning).
  • Multilingual: CC100, mC4 etc. multilingual corpora.

Key techniques

  • Data cleaning: dedup, quality filter (use models to score data), PII removal.
  • Data mix: how to mix code / encyclopedia / news / dialogue directly affects downstream capability.
  • Curriculum learning: easy-to-hard or vice versa.
  • Scaling Law: how to balance model size, data volume, compute (Chinchilla paper).