Skip to main content
ZICQ

Wiki Terminology

Data Leakage and Evaluation Contamination

Terminology
Aliases: benchmark contamination train-test leakage ·2026-10-07

Data Leakage and Evaluation Contamination

Data leakage occurs when training, preprocessing, or model selection uses information that should be unavailable during evaluation. It can inflate offline estimates of generalization. This entry concerns evaluation integrity; disclosure of private user data is a separate issue.

Common forms

  • Identical, near-duplicate, or closely related entity records cross training and test sets.
  • Scaling, imputation, or feature selection is fitted on the entire dataset before splitting.
  • A prediction uses information created after the event being predicted.
  • Benchmark questions, answers, or their rewrites enter training, causing evaluation contamination.

Example: support tickets

Imagine a ticket with an original question and two rewrites. Random row splitting can put related examples in both training and test sets. We would group by ticket ID and consider a temporal split when the deployment target is future tickets.

Suggested checks

Freeze test-set versions, check exact and approximate duplicates, trace synthetic-data sources, and inspect whether each feature is available at prediction time. Fit preprocessing on training data and transform validation and test data afterward.

Interpret cautiously

No exact string match does not establish the absence of leakage. With a black-box model, the full training corpus may be unknown. We suggest documenting the checks, coverage, and remaining uncertainty rather than inferring contamination from a high score alone.

Sources