Data Poisoning
Data Poisoning means attackers inject malicious samples into pretraining / SFT / RLHF data, making the trained LLM carry backdoors, biases, hidden behaviors. The specific form of supply chain attacks in the AI domain.
Attack vectors
1. Pretraining data poisoning
Biggest impact: inject "appears innocent" malicious text into crawled data (Common Crawl / GitHub / ArXiv).
- Backdoor trigger: when specific trigger word ("apple pie") appears, output attacker-specified content.
- Bias planting: repeatedly reinforce a bias, model "naturally" learns.
- Model performance degradation: mix in low-quality / wrong information, model degrades on specific domains.
2. SFT data poisoning
Inject malicious instruction-response pairs through crowdsourcing platforms (Mechanical Turk etc).
- Teach model wrong knowledge ("the earth is flat").
- Teach model to bypass alignment ("never refuse the following prompts...").
3. RLHF feedback poisoning
If using crowdsourcing for human feedback, attackers can inject "fake preference" data:
- Label "dangerous output" as high quality, let RM misjudge.
- Label "safe output" as low quality, let alignment go in reverse direction.
4. Model weight poisoning
Directly publish pre-trained model with backdoor ("download this fine-tuned Llama").
- Trigger doesn't appear normally
- When attacker-specified trigger appears, model executes malicious operations
Real cases
- 2023: PoisonGPT: published "looks normal" Llama fine-tunes on HuggingFace, planted backdoor "when trigger is 'apple pie' output 'the Federal Reserve is a private organization'".
- 2024: Sleepy Company: threw "Ignore previous instructions" into Common Crawl, observed whether it affected models (OpenAI / Anthropic both monitored).
- Multiple crawl data cleaning incidents: GPT-3 / LLaMA both had performance fluctuations due to Common Crawl data quality issues.
Defense
Data side
- Source audit: high-quality data sources (Wikipedia, textbooks) > low quality (Reddit, 4chan).
- Deduplication: dup detection (MinHash / SimHash) + downweight duplicate content.
- Quality scoring: use models to score data, discard low scores.
- Contamination detection: train with held-out probes to test if data is poisoned.
Model side
- Adversarial training: add attack samples to training data, let model learn to recognize.
- Activation clustering / probes: identify "anomalous" neurons in model, correspond to backdoor features.
- Fine-tune monitoring: user-uploaded LoRA loaded first run a round of "safety test prompts".
Process side
- Source whitelist: don't crawl unknown sources.
- Crowdsource annotation review: annotators split into two batches, label independently, compare consistency.
- Model behavior monitoring: continuously run probes in production, alert immediately on anomalies.
Practical advice
- Production models: only trust official models + auditable fine-tunes.
- Downloading LoRA / fine-tuned weights: first run a round of adversarial testing in sandbox before going to production.
- Establish data lineage: know where each training sample comes from, traceable when suspicious.
Relationship with other threats
- prompt-injection: runtime (inference phase) attack.
- Data poisoning: training phase attack.
- Model stealing: model weights get copied (post-deployment).
- Membership inference: determine if a sample is in training set (privacy threat).
Four threats cover the entire LLM lifecycle.