Pre-training
Pre-training is the first stage of LLM training: using massive unlabeled text (trillion-token scale), the model learns language knowledge and world knowledge through the self-supervised "next-token prediction" task.
Typical scale
- GPT-2: ~10B tokens
- LLaMA 1: ~1.4T tokens
- LLaMA 3: ~15T tokens
- Top industry models: tens of T tokens and up
Data sources
- Web pages: Common Crawl is the largest single source (60-80%), needs strict cleaning.
- Code: GitHub, StackOverflow (improve code ability).
- Books: high-quality long text (Project Gutenberg, Books3).
- Academic: arXiv papers, PubMed (improve scientific reasoning).
- Multilingual: CC100, mC4 etc. multilingual corpora.
Key techniques
- Data cleaning: dedup, quality filter (use models to score data), PII removal.
- Data mix: how to mix code / encyclopedia / news / dialogue directly affects downstream capability.
- Curriculum learning: easy-to-hard or vice versa.
- Scaling Law: how to balance model size, data volume, compute (Chinchilla paper).