arXiv Releases DEEPO: A Novel Reinforcement Learning Optimization Method to Reduce Hallucination in MLLMs
By Mr.Xu
Published:
Summary:arXiv has released a novel reinforcement learning optimization method called DEEPO (Dual-Entropy Enhanced Policy Optimization), designed to address hallucination issues in Multimodal Large Language Models (MLLMs). DEEPO employs a dual-stage enhancement strategy combining signal variance regularization and gradient preconditioning to address shortcomings in traditional RL methods when handling high-semantic-entropy queries and the vanishing of score-gradient norms. Experiments demonstrate that DE
Key Breakthroughs
-
Problem Context: Multimodal Large Language Models (MLLMs) excel in reasoning tasks but suffer from persistent hallucination issues. Traditional reinforcement learning (RL) methods often fail to handle high-semantic-entropy queries, leading to unanimous wrong sample groups and increased hallucination risk.
-
DEEPO Approach:
- Signal Variance Regularization: Semantic-entropy-triggered expert prefixes inject grounded continuations for high-uncertainty queries, restoring advantage variance.
- Gradient Preconditioning: Advantage-sign-aware Renyi preconditioning addresses logit-level saturation, allowing corrections to reach confident errors in the operational confidence regime.
-
Experimental Results:
- On VideoMMMU (the most complex long-horizon task), DEEPO outperformed GRPO by 4.0% (95% CI [1.1, 6.9]).
- In other tasks, DEEPO demonstrated stable performance improvements while maintaining model accuracy and training stability.
Technical Highlights
- Dual-Stage Enhancement Strategy: Combines signal variance regularization and gradient preconditioning to address the shortcomings of traditional RL methods.
- Semantic-Entropy Trigger Mechanism: Provides direct supervision for high-uncertainty queries through expert prefixes.
- Advantage-Sign-Aware Preconditioning: Effectively addresses logit-level saturation, enabling more precise corrections.
Industry Impact
- MLLM Optimization: Offers a new technical path for addressing hallucination issues, enhancing the reliability of models in practical applications.
- RL Application Expansion: Demonstrates the potential of RL in optimizing complex tasks, promoting further development of RL in the AI field.
Developer Recommendations
- Model Optimization: Developers are advised to consider applying DEEPO to optimize their existing MLLMs for improved performance.
- Experimental Validation: When applying DEEPO, thorough experimental validation is recommended to ensure its effectiveness in specific tasks.
- Continuous Attention: As RL technology evolves, developers should stay updated on the latest research in the field.
— END —Source: ArXiv AI (cs.AI) (2026-09-25)
Tags: #DEEPO #Reinforcement Learning #Multimodal Large Language Models #Hallucination Reduction #arXiv
Community Comments