Hugging Face Study Reveals Over-Editing in Code Repair, Proposes Minimal Editing Optimization Strategies
By Mr.Xu
Published: · 4 views
Summary:Hugging Face's research team investigates the issue of over-editing in code repair by large language models (LLMs). They propose an evaluation framework that injects AST-level corruptions into reference solutions to quantify excessive edits. The study reveals that even strong models like GPT-5.5 exhibit over-editing tendencies, leading to unnecessary complexity and reduced efficiency. By introducing a preservation instruction, the average excess Levenshtein distance is reduced, cognitive complex
Background and Problem
Large language models (LLMs) are increasingly used for code repair, but ensuring correctness alone is insufficient. Useful repairs should also minimize changes, maintain reviewability, and remain faithful to the original implementation. However, existing research indicates that many models suffer from over-editing, where unnecessary modifications are made during bug fixes, leading to increased code complexity and reduced reviewability.
Methodology and Findings
The research team developed an evaluation framework that injects controlled AST-level corruptions into reference solutions, assigning each repair task a known minimal patch. Testing across frontier LLMs revealed that over-editing is widespread, even among strong models like GPT-5.5.
The study also found that introducing a preservation instruction significantly reduces over-editing, lowering the average excess Levenshtein distance from 0.195 to 0.131, reducing cognitive complexity by 26.6%, and improving Pass@1 by 2.3 points. These findings suggest that optimizing instructions can effectively improve model editing behavior.
Optimization Strategies and Results
To further optimize minimal editing behavior, the team explored the possibility of directly learning minimal edits during post-training. The results showed that supervised fine-tuning tends to overfit to seen corruption patterns, whereas reinforcement learning offers the best trade-off for out-of-domain edit fidelity and performance retention.
Conclusion and Impact
This study positions edit fidelity as a distinct axis of code-repair quality and demonstrates its measurability and learnability. The findings provide new optimization directions for AI-driven code repair, particularly in scenarios requiring high fidelity repairs.
Recommendations for Developers
- Optimize Instruction Design: When applying LLMs for code repair, consider designing preservation instructions to reduce over-editing.
- Explore Reinforcement Learning: For tasks requiring high fidelity repairs, explore the use of reinforcement learning for post-training optimization.
- Focus on Edit Fidelity: When evaluating code repair models, prioritize edit fidelity as a key metric.
— END —Source: Hugging Face Daily Papers (2026-09-03)
Tags: #LLMs & Foundation Models #Code Repair #Reinforcement Learning #Hugging Face #AI Research
Community Comments