ZICQ
中 Log in / Sign up
Newsroom Agentic #DeReAct #AI Agents #Modular Architecture #Task Completion Control #ReAct Improvement

DeReAct Architecture Released: A New Approach to Enhance AI Agent Reliability in Decision-Making and Task Completion

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:arXiv introduces DeReAct, a novel AI agent architecture designed to address the limitations of traditional ReAct-based agents in action authorization and task completion control. By incorporating two distinct gating policies—a Critic for validating actions and a Context Manager for certifying task completion—DeReAct significantly improves the performance of weaker models. In benchmarks like GAIA and SWE-bench Verified, DeReAct boosts Pass@1 scores by 6.5-7.0 and 4.2-5.2 percentage points for Qwe


1. Introduction

In the realm of AI agents, traditional ReAct methods rely on a single large language model (LLM) policy to propose actions, interact with the environment, and decide when a task is complete. This approach suffers from difficulties in independently enforcing action authorization and task completion control, leading to error propagation and unsupported completion claims that terminate execution prematurely.

2. DeReAct Architecture

DeReAct addresses these issues through a modular design, featuring two key innovations:

  • Critic Module: Validates proposed actions before execution to ensure their reasonableness and safety.
  • Context Manager Module: Reconstructs an environment-supported State and certifies task completion.

This design decouples action validation and task completion control, enabling AI agents to operate more reliably in complex tasks.

3. Experimental Results

In benchmarks like GAIA and SWE-bench Verified, DeReAct shows the most significant improvements for weaker models:

  • Qwen3-Coder-480B: Achieves a 6.5-7.0 percentage point increase in Pass@1.
  • Claude Sonnet 4.5: Achieves a 4.2-5.2 percentage point increase in Pass@1.

As model capabilities increase, the gains diminish. For instance, with Claude Opus 4.5, Pass@1 remains comparable to ReAct, but DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that it trades earlier termination for stronger grounding.

4. Technical Highlights

  • Modular Design: Separates action validation and task completion control to enhance system reliability.
  • External Gating Policies: Utilizes Critic and Context Manager for fine-grained control over actions and task completion.
  • Performance Improvement: Significantly boosts the performance of weaker models while maintaining the robustness of stronger ones.

5. Industry Impact

DeReAct's release offers new avenues for enhancing AI agent reliability in complex tasks, particularly in high-stakes applications such as autonomous driving, medical diagnostics, and industrial automation. The modular design also provides a framework for future optimization of AI agent architectures.

6. Developer Recommendations

  • Emphasize Modular Design: Consider separating action validation and task completion control in AI agent design to improve system reliability.
  • Test Across Model Capabilities: Evaluate DeReAct's performance across models of varying capabilities to assess its impact on model performance.
  • Focus on Task Completion Certification: Adopt the Context Manager's design principles to ensure accurate and reliable task completion.

Source: ArXiv AI (cs.AI) (2026-10-05)

— END —

Tags: #DeReAct #AI Agents #Modular Architecture #Task Completion Control #ReAct Improvement

Community Comments

Loading live comments and annotations…