ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #DEEPO #Reinforcement Learning #Multimodal Large Language Models #Hallucination Reduction #arXiv

arXiv Releases DEEPO: A Novel Reinforcement Learning Optimization Method to Reduce Hallucination in MLLMs

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:arXiv has released a novel reinforcement learning optimization method called DEEPO (Dual-Entropy Enhanced Policy Optimization), designed to address hallucination issues in Multimodal Large Language Models (MLLMs). DEEPO employs a dual-stage enhancement strategy combining signal variance regularization and gradient preconditioning to address shortcomings in traditional RL methods when handling high-semantic-entropy queries and the vanishing of score-gradient norms. Experiments demonstrate that DE


Key Breakthroughs

  • Problem Context: Multimodal Large Language Models (MLLMs) excel in reasoning tasks but suffer from persistent hallucination issues. Traditional reinforcement learning (RL) methods often fail to handle high-semantic-entropy queries, leading to unanimous wrong sample groups and increased hallucination risk.

  • DEEPO Approach:

    1. Signal Variance Regularization: Semantic-entropy-triggered expert prefixes inject grounded continuations for high-uncertainty queries, restoring advantage variance.
    2. Gradient Preconditioning: Advantage-sign-aware Renyi preconditioning addresses logit-level saturation, allowing corrections to reach confident errors in the operational confidence regime.
  • Experimental Results:

    • On VideoMMMU (the most complex long-horizon task), DEEPO outperformed GRPO by 4.0% (95% CI [1.1, 6.9]).
    • In other tasks, DEEPO demonstrated stable performance improvements while maintaining model accuracy and training stability.

Technical Highlights

  • Dual-Stage Enhancement Strategy: Combines signal variance regularization and gradient preconditioning to address the shortcomings of traditional RL methods.
  • Semantic-Entropy Trigger Mechanism: Provides direct supervision for high-uncertainty queries through expert prefixes.
  • Advantage-Sign-Aware Preconditioning: Effectively addresses logit-level saturation, enabling more precise corrections.

Industry Impact

  • MLLM Optimization: Offers a new technical path for addressing hallucination issues, enhancing the reliability of models in practical applications.
  • RL Application Expansion: Demonstrates the potential of RL in optimizing complex tasks, promoting further development of RL in the AI field.

Developer Recommendations

  • Model Optimization: Developers are advised to consider applying DEEPO to optimize their existing MLLMs for improved performance.
  • Experimental Validation: When applying DEEPO, thorough experimental validation is recommended to ensure its effectiveness in specific tasks.
  • Continuous Attention: As RL technology evolves, developers should stay updated on the latest research in the field.

Source: ArXiv AI (cs.AI) (2026-09-25)

— END —

Tags: #DEEPO #Reinforcement Learning #Multimodal Large Language Models #Hallucination Reduction #arXiv

Community Comments

Loading live comments and annotations…