Data Leakage and Evaluation Contamination
Data leakage occurs when training, preprocessing, or model selection uses information that should be unavailable during evaluation. It can inflate offline estimates of generalization. This entry concerns evaluation integrity; disclosure of private user data is a separate issue.
Common forms
- Identical, near-duplicate, or closely related entity records cross training and test sets.
- Scaling, imputation, or feature selection is fitted on the entire dataset before splitting.
- A prediction uses information created after the event being predicted.
- Benchmark questions, answers, or their rewrites enter training, causing evaluation contamination.
Example: support tickets
Imagine a ticket with an original question and two rewrites. Random row splitting can put related examples in both training and test sets. We would group by ticket ID and consider a temporal split when the deployment target is future tickets.
Suggested checks
Freeze test-set versions, check exact and approximate duplicates, trace synthetic-data sources, and inspect whether each feature is available at prediction time. Fit preprocessing on training data and transform validation and test data afterward.
Interpret cautiously
No exact string match does not establish the absence of leakage. With a black-box model, the full training corpus may be unknown. We suggest documenting the checks, coverage, and remaining uncertainty rather than inferring contamination from a high score alone.
Sources
- scikit-learn: Data leakage — definition and preprocessing leakage.
- scikit-learn: Cross-validation — grouped and temporal evaluation.
- Deduplicating Training Data Makes Language Models Better — training/evaluation overlap.