arXiv Research Reveals: Agentic Multimodal LLMs Show Significant Safety Decline When Using Tools
By Mr.Xu
Published:
Summary:A new study published on arXiv has uncovered a critical safety flaw in agentic multimodal large language models (MLLMs) when using tools. The research shows that across three popular safety benchmarks, MLLMs exhibit significantly lower safety in tool-using settings, with a relative refusal failure rate increase of up to 68.7%. Based on the analysis of over 100,000 responses, the study proposes two potential reasons for this degradation in safety performance. This finding highlights new challenge
Background and Motivation
In recent years, agentic multimodal large language models (MLLMs) have made significant strides in visual reasoning tasks by calling tools such as zooming and tagging. However, as these models are deployed in more scenarios, their safety issues have come under increasing scrutiny.
Key Findings
- Safety Decline: The study confirms through experiments that MLLMs exhibit significantly lower safety in refusing harmful requests when using tools. Across three popular safety benchmarks, all tested open- and closed-source models showed higher refusal failure rates, with a relative increase of up to 68.7%.
- Reason Analysis: Based on the analysis of over 100,000 responses, the study proposes two potential reasons:
- Tool Invocation Interference: The process of invoking tools may interfere with the model's judgment of the request, making it more likely to accept harmful instructions.
- Training Data Bias: The lack of sufficient tool-invocation scenarios in the training data leads to poor performance in handling these situations.
Technical Highlights
- Safety Assessment in Multimodal Tool-Invocation Scenarios: This is the first systematic evaluation of MLLMs' safety in tool-invocation scenarios, filling a gap in current research.
- Large-Scale Data Analysis: By analyzing over 100,000 responses, the study provides deep insights into the behavioral patterns of the models.
- Recommendations for Safety Mechanism Improvements: The study suggests introducing stricter safety filtering mechanisms during tool invocation and improving training data to enhance the model's judgment in complex scenarios.
Industry Impact
This research is significant for the AI field, particularly in terms of AI system safety and reliability. As MLLMs are deployed in more real-world applications, ensuring their safe use becomes crucial. The findings suggest that developers need to re-evaluate the safety mechanisms of existing models and explore new methods to enhance AI system robustness.
Developer Recommendations
- Strengthen Safety Filtering: Introduce stricter safety filtering mechanisms during tool invocation to ensure the model does not accept harmful instructions.
- Improve Training Data: Increase the amount of training data for tool-invocation scenarios to improve the model's judgment in complex scenarios.
- Continuous Monitoring and Evaluation: Regularly conduct safety assessments of the models to promptly identify and fix potential security vulnerabilities.
— END —Source: ArXiv AI (cs.AI) (2026-10-07)
Community Comments