Hugging Face Research: Exploring Synthetic Data Attribution and Training Data Selection Methods
By Mr.Xu
Published:
Summary:Hugging Face's research team has published a study on synthetic data attribution and training data selection. The study examines the reliability of attributing generated passages to their sources and whether this information aids in selecting better training data. The results show that generator attribution is highly accurate (98.7%) for original passages but drops significantly after paraphrasing (53.1%) and style rewriting (29.0%). Additionally, the research compares two methods for selecting
Background and Motivation
With the widespread use of generative AI models, repeated training on model-generated data can degrade the performance of subsequent models. To address this, researchers propose using provenance information to select generated data. However, the reliability of provenance information and its effectiveness in selecting training data need further validation.
Methodology and Experiments
The research team used financial risk text as experimental data. They first tested the accuracy of generator attribution for original passages and then repeated the test after rewriting the text (including paraphrasing and style rewriting). The results showed that the attribution accuracy for original passages was 98.7%, but it dropped to 53.1% after paraphrasing and 29.0% after style rewriting.
For data selection, the study compared two methods:
- Source-based selection: Using provenance information provided by the generator to select data.
- Reference model score-based selection: Using a separate reference model to score the generated data and selecting data with higher scores.
Key Findings
- Limited reliability of provenance information: While the attribution accuracy for original passages is high, it significantly decreases after rewriting, indicating limitations in handling complex texts.
- No significant difference between selection methods: During three rounds of generation and retraining, the two selection methods chose different data samples but did not show a significant difference in the degradation of the resulting models.
Conclusions and Implications
The study concludes that identifying data provenance and selecting useful training data are separate challenges. Current methods, including provenance scoring and reference model scoring, may not be sufficient for future recursive training behaviors. This poses new challenges for researchers to develop more reliable data selection methods to support recursive training.
Recommendations for Developers
- Use generated data cautiously: When training models, use generated data cautiously to avoid degradation in model performance due to data quality issues.
- Combine multiple selection methods: Developers can try combining multiple data selection methods to improve the quality of training data.
- Be aware of provenance limitations: When relying on provenance information, be aware of its limitations and consider other supplementary methods.
— END —Source: Hugging Face Daily Papers (2026-09-30)
Tags: #Hugging Face #Data Attribution #Training Data Selection #Generative AI #Model Training
Community Comments