ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Diffusion Models #Model Security #Bias Injection #AI Safety

Hugging Face Proposes Novel Method: Targeted Bias Injection via Noise Trajectory in Diffusion Language Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face researchers introduce a novel attack method that injects targeted bias into diffusion language models (dLLMs) via noise trajectory manipulation. By exploiting the dLLM's iterative denoising process, the method uses a proportional-integral (PI) controller to adaptively steer the model toward a chosen demographic answer. Experiments show significant increases in targeted preference, such as raising LLaDA-8B-Instruct's preference for a targeted group from 1.8% to 16.7% and increasing s


Background and Motivation

Diffusion language models (dLLMs) generate text through iterative denoising, re-predicting masked positions multiple times before committing to an answer. Unlike autoregressive decoders, dLLMs expose the distribution of the answer at every denoising step before the final commitment, which provides potential attackers with opportunities to manipulate the model's output.

Method and Innovation

The study proposes a novel method for injecting targeted bias into dLLMs via noise trajectory manipulation, with the following steps:

  1. Attack Mechanism: A proportional-integral (PI) controller is used to track the probability of the target answer in real-time and adaptively adjust the intervention strength.
  2. Experimental Validation: On the LLaDA-8B-Instruct model, the attack increases the preference for the targeted group from 1.8% to 16.7% on ambiguous BBQ questions (where the correct answer is abstention), which is more than three times the strongest fixed-strength intervention baseline. On the SocialStigmaQA dataset, the selection of stigmatizing answers rises from 17.6% to 58.1%.
  3. Extended Applications: For other demographic targets, the attack can shift answers by up to 37 percentage points, with each attack taking about 40 minutes on a single GPU.

Technical Highlights

  • Novel Attack Path: This is the first study to reveal the control channel in the denoising trajectory of dLLMs, providing a new perspective for model security research.
  • Efficient Intervention Strategy: The PI controller can dynamically adjust the intervention strength based on the probability of the target answer, significantly enhancing the attack effect.
  • Wide Applicability: The method can be customized for different demographic targets, demonstrating its strong generalization ability.

Industry Impact and Recommendations

  1. Importance of Security Audits: The study emphasizes the necessity of more comprehensive bias audits for dLLMs, considering not only the static model but also the dynamic behavior of the serving stack.
  2. Developer Recommendations: When deploying dLLMs, additional security mechanisms, such as real-time monitoring and intervention, should be introduced to prevent potential bias injection attacks.
  3. Future Research Directions: Further exploration of more complex attack paths and defense strategies is needed to enhance the robustness and security of dLLMs.

Conclusion

This research exposes potential security vulnerabilities in dLLMs during the denoising process and proposes an effective method for injecting targeted bias via noise trajectory manipulation. The findings have significant implications for the security and reliability of AI models, urging the AI community to strengthen bias audits and security measures for dLLMs.


Source: Hugging Face Daily Papers (2026-10-05)

— END —

Tags: #Hugging Face #Diffusion Models #Model Security #Bias Injection #AI Safety

Community Comments

Loading live comments and annotations…