ArXiv Introduces RACER: A Novel Framework for Backdoor Repair in Multimodal Large Language Models
By Mr.Xu
Published: · 2 views
Summary:ArXiv researchers introduce RACER, a model-level repair framework designed to eliminate latent backdoors in Multimodal Large Language Models (MLLMs). RACER operates by identifying abnormal layer-to-layer inconsistencies in internal representations, focusing on token regions that encode trigger features. The framework decomposes fused representations into visual and textual token regions, normalizes their inconsistencies separately, and recomposes them using modality-aware weights. RACER requires
Overview
Multimodal Large Language Models (MLLMs) are becoming increasingly prevalent in user-facing applications. However, they inherit backdoor risks from the construction pipelines, posing significant security concerns. Existing model-level backdoor removal methods, primarily designed for traditional classifiers, show limited effectiveness on MLLMs. In response, the ArXiv team introduces RACER, a novel model-level repair framework aimed at eliminating latent backdoors in MLLMs at their source.
Key Features
-
Layer-wise Inconsistency Anomaly Analysis: RACER identifies abnormal layer-to-layer inconsistencies in internal representations, attributing them to the token regions encoding trigger features.
-
Modality-aware Normalization: The framework decomposes the fused representation into visual and textual token regions and normalizes their inconsistencies separately.
-
Modality-aware Weight Recomposition over Deep-layer Window: RACER uses modality-aware weights to recompose the token regions over a deep-layer window, better capturing localized backdoor-induced anomalies.
-
Adversarial Fine-tuning and Worst-case Perturbation Synthesis: RACER employs a min-max optimization approach to synthesize worst-case perturbations and perform adversarial fine-tuning, repairing the model by suppressing the deep representational directional shifts on which backdoor behaviors rely.
Experimental Results
Evaluations on three open-source MLLMs across 36 backdoor settings, including image, text, and multimodal triggers, demonstrate that RACER reduces the average Attack Success Rate (ASR) to 1.1%, achieving 0% in 32 settings. Moreover, RACER preserves the clean-task utility of the model, ensuring that the model's performance on legitimate tasks remains unaffected.
Industry Impact and Developer Recommendations
- Enhanced Model Security: RACER provides an effective method for backdoor repair in MLLMs, contributing to the security and reliability of models deployed in critical applications.
- No Trigger Knowledge Required: The framework does not require knowledge of the trigger or attack objective, simplifying the repair process.
- Resource Efficiency: RACER requires only 100 clean samples, making it suitable for resource-constrained environments.
Developers are encouraged to integrate RACER with other defense mechanisms to build more robust security frameworks for MLLMs. Additionally, it is recommended to perform thorough backdoor detection and repair before deploying models to ensure their security and reliability.
— END —Source: ArXiv cs.AI (2026-08-25)
Tags: #RACER #Backdoor Repair #Multimodal Large Language Models #Model Security
Community Comments