ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #ArXiv #Long-Document QA #VQA #Memory Modeling #GRPO

ArXiv Introduces AWM Framework: Enhancing Memory Quality and Reasoning Accuracy for Long-Document Visual Question Answer

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The ArXiv team introduces the Answerable Working Memory (AWM) framework to enhance memory quality and reasoning accuracy in long-document visual question answering (VQA) tasks. AWM treats terminal working memory as an answerable evidence artifact and incorporates this signal into the GRPO reward while preserving final-answer priority. Experiments on MMLongBench-Doc and LongDocURL datasets demonstrate that AWM-GRPO significantly improves final-answer accuracy and reduces the memory-missing-correc


Core Breakthrough

The ArXiv team introduces the Answerable Working Memory (AWM) framework to address the challenge of insufficient memory quality in long-document visual question answering (VQA) tasks. The key innovations include:

  • Memory-Only Answerability Diagnostic: This evaluates whether the terminal working memory alone can support answering the question, identifying gaps in memory quality.
  • AWM-GRPO Mechanism: This integrates the memory-only answerability signal into the GRPO reward, maintaining the priority of the final answer while placing higher weight on memory quality.

Technical Highlights

  1. Challenges in Memory Modeling for Long-Document VQA: Existing methods primarily focus on the correctness of the final answer and access to evidence pages, but overlook the quality of working memory. AWM fills this gap by introducing a new diagnostic approach.
  2. Improvement to GRPO Reward Mechanism: AWM-GRPO builds on the GRPO mechanism by adding a reward for memory quality, enabling the model to better utilize information in working memory when generating answers.
  3. Significant Experimental Results: Experiments on MMLongBench-Doc and LongDocURL datasets show that AWM-GRPO outperforms existing methods in both final answer accuracy and memory quality. For instance, on MMLongBench-Doc, the final answer accuracy improved by 8.1 percentage points.

Industry Impact

The AWM framework offers a new approach to memory modeling in long-document question answering systems, significantly enhancing the reasoning capabilities and answer quality of models, especially in handling complex documents and multi-step reasoning tasks. This is particularly valuable for AI applications that require processing long documents, such as legal document analysis, medical record processing, and academic literature review.

Developer Recommendations

  • Expand Application Scenarios: Developers can apply the AWM framework to AI applications that involve processing long documents and multi-step reasoning, such as legal consulting systems and medical diagnosis assistants.
  • Model Optimization: By combining the AWM framework, developers can further optimize the memory modeling mechanism of existing models, improving their performance in complex tasks.
  • Cross-Domain Application: The AWM framework is not limited to visual question answering tasks and can be extended to other areas that require high-quality memory modeling, such as natural language inference and dialogue systems.

Source: ArXiv cs.CL (2026-08-26)

— END —

Tags: #ArXiv #Long-Document QA #VQA #Memory Modeling #GRPO

Community Comments

Loading live comments and annotations…