Skip to main content
ZICQ

Wiki Terminology

Data Cleaning

Terminology
Aliases: deduplication data preprocessing ·2026-10-07

Data Cleaning

Data cleaning identifies and handles duplicates, malformed records, missing values, and content unsuitable for a task. Its purpose is to produce traceable inputs for training or retrieval.

A suggested workflow

  1. Preserve a raw snapshot, stable IDs, provenance, and transformation versions.
  2. Validate encoding, field types, and missing values; separate page navigation from useful text.
  3. Define task-specific filters and inspect samples of retained and rejected records.
  4. Use exact deduplication for identical records and approximate checks for near-duplicates.
  5. Check overlap between splits and record removal counts and reasons.

This is an engineering proposal, not a universal ordering. Fit statistical transformations on training data only; see data-leakage.

What deduplication establishes

Lee and colleagues found reduced verbatim memorization and training/evaluation overlap in their studied language-model datasets. Their results do not imply an identical benefit for every task.

Example: a help center

Suppose three URLs contain the same help page. We would keep one canonical text and retain a source mapping. If an updated procedure replaces an older one, track versions instead of combining contradictory instructions. Avoid mechanically deleting code, negation, numbers, or table structure.

Check the outcome

Our suggested review compares duplicate rates, parsing failures, task coverage, and sampled errors. Fewer records alone do not establish better quality; aggressive filters can remove rare languages and edge cases.

Sources