Skip to main content
ZICQ

Wiki Security

Data Poisoning

Security
Aliases: data poisoning training data attack ·2026-09-14

Data Poisoning

Data Poisoning means attackers inject malicious samples into pretraining / SFT / RLHF data, making the trained LLM carry backdoors, biases, hidden behaviors. The specific form of supply chain attacks in the AI domain.

Attack vectors

1. Pretraining data poisoning

Biggest impact: inject "appears innocent" malicious text into crawled data (Common Crawl / GitHub / ArXiv).

  • Backdoor trigger: when specific trigger word ("apple pie") appears, output attacker-specified content.
  • Bias planting: repeatedly reinforce a bias, model "naturally" learns.
  • Model performance degradation: mix in low-quality / wrong information, model degrades on specific domains.

2. SFT data poisoning

Inject malicious instruction-response pairs through crowdsourcing platforms (Mechanical Turk etc).

  • Teach model wrong knowledge ("the earth is flat").
  • Teach model to bypass alignment ("never refuse the following prompts...").

3. RLHF feedback poisoning

If using crowdsourcing for human feedback, attackers can inject "fake preference" data:

  • Label "dangerous output" as high quality, let RM misjudge.
  • Label "safe output" as low quality, let alignment go in reverse direction.

4. Model weight poisoning

Directly publish pre-trained model with backdoor ("download this fine-tuned Llama").

  • Trigger doesn't appear normally
  • When attacker-specified trigger appears, model executes malicious operations

Real cases

  • 2023: PoisonGPT: published "looks normal" Llama fine-tunes on HuggingFace, planted backdoor "when trigger is 'apple pie' output 'the Federal Reserve is a private organization'".
  • 2024: Sleepy Company: threw "Ignore previous instructions" into Common Crawl, observed whether it affected models (OpenAI / Anthropic both monitored).
  • Multiple crawl data cleaning incidents: GPT-3 / LLaMA both had performance fluctuations due to Common Crawl data quality issues.

Defense

Data side

  • Source audit: high-quality data sources (Wikipedia, textbooks) > low quality (Reddit, 4chan).
  • Deduplication: dup detection (MinHash / SimHash) + downweight duplicate content.
  • Quality scoring: use models to score data, discard low scores.
  • Contamination detection: train with held-out probes to test if data is poisoned.

Model side

  • Adversarial training: add attack samples to training data, let model learn to recognize.
  • Activation clustering / probes: identify "anomalous" neurons in model, correspond to backdoor features.
  • Fine-tune monitoring: user-uploaded LoRA loaded first run a round of "safety test prompts".

Process side

  • Source whitelist: don't crawl unknown sources.
  • Crowdsource annotation review: annotators split into two batches, label independently, compare consistency.
  • Model behavior monitoring: continuously run probes in production, alert immediately on anomalies.

Practical advice

  • Production models: only trust official models + auditable fine-tunes.
  • Downloading LoRA / fine-tuned weights: first run a round of adversarial testing in sandbox before going to production.
  • Establish data lineage: know where each training sample comes from, traceable when suspicious.

Relationship with other threats

  • prompt-injection: runtime (inference phase) attack.
  • Data poisoning: training phase attack.
  • Model stealing: model weights get copied (post-deployment).
  • Membership inference: determine if a sample is in training set (privacy threat).

Four threats cover the entire LLM lifecycle.