ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Backdoor Attacks #LLM Agents #AI Security #UIUC Kang Lab #PersistBD

UIUC Kang Lab Releases PersistBD: Unveiling Backdoor Attack Persistency Risks in LLM Agents

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The UIUC Kang Lab research team has released a study on the persistence of backdoor attacks in LLM agents, introducing a technique called PersistBD. The research investigates whether backdoors in models supplied by attackers can survive through benign post-training processes such as supervised fine-tuning (SFT) and reinforcement learning (RL). The findings reveal that while SFT reduces attack success rates, subsequent RL often preserves or even amplifies residual backdoor behaviors. PersistBD re


Background and Motivation

In the field of backdoor attacks, adversaries implant hidden behaviors into models that produce malicious outputs when specific input patterns are present. This study focuses on whether such backdoors can survive through benign post-training processes, including supervised fine-tuning (SFT) and task-level reinforcement learning (RL), when developers adapt third-party models.

Key Findings

  1. Persistence of Backdoors: The research shows that while SFT significantly reduces attack success rates, subsequent RL often preserves or even amplifies residual backdoor behaviors.
  2. Key Factors: The persistence of backdoors is primarily influenced by two factors: the initial strength of the backdoor and its gradient compatibility with benign training.
  3. PersistBD Proposal: To address this issue, the team proposes PersistBD, a technique that refines backdoored models before release to enhance their persistence through benign training.

Experimental Results

Experiments on the Qwen2.5-Coder-7B model demonstrate that PersistBD increases attack success rates from 20% to 74% after SFT and to 76% after SFT-RL, while maintaining comparable performance on benign tasks.

Industry Impact and Recommendations

  1. Supply Chain Risks: Backdoor attacks pose significant risks to the AI development supply chain, and developers should strengthen the detection and validation of third-party models.
  2. Technical Needs: Current backdoor detection and mitigation techniques need further improvement to address increasingly sophisticated security threats.
  3. Developer Recommendations: AI developers are advised to adopt stricter security assessment processes when using third-party models and consider using new technologies like PersistBD to enhance backdoor detection and mitigation capabilities.

Technical Highlights

  • PersistBD Technology: Optimizes the structure of backdoored models to maintain high attack success rates after benign training.
  • Experimental Validation: Validates the effectiveness of PersistBD through multiple benchmarks, demonstrating its persistence through complex training processes.

Conclusion

This study highlights the persistence risks of backdoor attacks in LLM agents and emphasizes the importance of strengthening security measures when adopting third-party models. PersistBD provides a new technical path for backdoor detection and mitigation, advancing the field of AI security.


Source: Hugging Face Daily Papers (2026-10-05)

— END —

Tags: #Backdoor Attacks #LLM Agents #AI Security #UIUC Kang Lab #PersistBD

Community Comments

Loading live comments and annotations…