Data Cleaning
Data cleaning identifies and handles duplicates, malformed records, missing values, and content unsuitable for a task. Its purpose is to produce traceable inputs for training or retrieval.
A suggested workflow
- Preserve a raw snapshot, stable IDs, provenance, and transformation versions.
- Validate encoding, field types, and missing values; separate page navigation from useful text.
- Define task-specific filters and inspect samples of retained and rejected records.
- Use exact deduplication for identical records and approximate checks for near-duplicates.
- Check overlap between splits and record removal counts and reasons.
This is an engineering proposal, not a universal ordering. Fit statistical transformations on training data only; see data-leakage.
What deduplication establishes
Lee and colleagues found reduced verbatim memorization and training/evaluation overlap in their studied language-model datasets. Their results do not imply an identical benefit for every task.
Example: a help center
Suppose three URLs contain the same help page. We would keep one canonical text and retain a source mapping. If an updated procedure replaces an older one, track versions instead of combining contradictory instructions. Avoid mechanically deleting code, negation, numbers, or table structure.
Check the outcome
Our suggested review compares duplicate rates, parsing failures, task coverage, and sampled errors. Fewer records alone do not establish better quality; aggressive filters can remove rare languages and edge cases.
Sources
- Deduplicating Training Data Makes Language Models Better — experimental findings and scope.
- Hugging Face Datasets: Process — mapping, filtering, and transformations.
- scikit-learn: Common pitfalls — preprocessing and leakage.