ZICQ
中 Log in / Sign up
Newsroom Research & Papers #ArXiv #Language Models #Optimiser #Forgetting Mechanism #Fine-Tuning

ArXiv Research: How Optimiser Memory Causes Forgetting in Language Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:ArXiv has published a study investigating the mechanisms of forgetting in language models during fine-tuning, identifying optimiser memory as a key factor. The research shows that momentum optimisers accumulate historical gradients, which can counteract the current gradient and weaken the model's retention of previously learned knowledge. By decomposing Adam updates, the study reveals the adversarial relationship between accumulated history and current gradients and suggests that interventions l


Background and Motivation

During the fine-tuning of language models, models sometimes forget previously learned knowledge even when the current gradient update aims to preserve it. This phenomenon is known as catastrophic forgetting. ArXiv's latest research delves into how optimiser memory contributes to this forgetting.

Key Findings

  1. Impact of Optimiser Memory: The study identifies that momentum optimisers (e.g., Adam) accumulate historical gradients during updates, and these gradients can counteract the current gradient, weakening the model's retention of old knowledge.
  2. Decomposition of Forgetting Mechanisms: By decomposing the old-task loss into confusion among answers and leakage outside the answer set, the research shows that accumulated history tends to promote leakage, while the current gradient resists it.
  3. Time Dependency: The influence of historical gradients changes over time, with newer gradients tending to protect old answers and older gradients mainly producing harmful effects.
  4. Effectiveness of Interventions: By resetting momentum and matching the initial update norm, the study demonstrates that stored history does affect retention, and longer Adam continuations with such interventions lead to lower final old-task loss, primarily through recovered answer mass.

Technical Highlights

  • Quantification of Optimiser Memory: The study is the first to systematically quantify the impact of historical gradients in momentum optimisers on forgetting.
  • Decomposition of Adam Updates: A new method is proposed to decompose Adam updates into the effects of historical and current gradients, providing a new perspective on understanding optimiser behavior.
  • Validation of Intervention Strategies: Experiments validate the effectiveness of interventions like resetting momentum in reducing forgetting.

Industry Impact and Developer Recommendations

  • Directions for Optimiser Improvement: The findings suggest new directions for optimiser design, such as adjusting the momentum accumulation mechanism to reduce forgetting.
  • Adjusting Model Training Strategies: Developers can experiment with introducing momentum reset mechanisms during training to improve the model's retention of old knowledge.
  • Designing Long-Term Memory Models: For tasks requiring long-term memory, it may be beneficial to design specialised optimisers or training strategies that better balance the influence of current updates and historical gradients.

Conclusion

This research reveals the critical role of optimiser memory in language model forgetting and provides new technical paths for reducing forgetting. By gaining a deeper understanding of optimiser behavior, AI researchers can develop more efficient and reliable model training methods.


Source: ArXiv NLP/LLM (cs.CL) (2026-10-06)

— END —

Tags: #ArXiv #Language Models #Optimiser #Forgetting Mechanism #Fine-Tuning

Community Comments

Loading live comments and annotations…