arXiv Releases XiangqiBench: A Benchmark for Closed-Loop Evaluation of LLM Agents in Chinese Chess
By Mr.Xu
Published:
Summary:arXiv has introduced XiangqiBench, a benchmark designed to evaluate the closed-loop performance of LLM agents in Chinese chess. The benchmark utilizes 119 endgame scenarios with forced mates, challenging agents to deliver checkmate against an engine defender. XiangqiBench employs an interactive REPL interface to separate real moves, state queries, and forward simulation, recording 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. The study reveals significant g
Key Breakthroughs
The release of XiangqiBench marks a significant advancement in the evaluation of LLM agents, with the following key features:
- Closed-loop Evaluation: Unlike traditional methods that focus on static move selection, XiangqiBench emphasizes the agent's closed-loop performance in adversarial environments.
- Multi-turn Trajectory Recording: By recording 8,568 multi-turn trajectories, XiangqiBench provides rich experimental data for analyzing agent performance in diverse scenarios.
- Interactive REPL Interface: This interface allows agents to perform state queries and forward simulation, more accurately simulating real-world applications.
Technical Highlights
- Forced Mate Endgames: XiangqiBench includes 119 endgame scenarios with forced mates, providing clear victory conditions for agents.
- Two Observation Protocols: Agents are tested under two different observation protocols to evaluate their performance under varying conditions.
- Multi-model Comparison: 12 frontier LLMs participated in the tests, offering ample comparative data.
Industry Impact
The release of XiangqiBench sets a new standard for evaluating LLM agents in closed-loop tasks. Its findings indicate that current evaluation methods may overestimate an agent's actual capabilities, particularly in adversarial environments. This will prompt AI researchers and developers to reconsider their evaluation methods and push for more reliable agent development.
Recommendations for Developers
- Focus on Closed-loop Evaluation: In agent development, pay more attention to their performance in closed-loop tasks, not just static move selection.
- Leverage Multi-turn Data: Use the multi-turn trajectory data provided by XiangqiBench to analyze agent performance in different scenarios and optimize accordingly.
- Incorporate Interactive Interfaces: Introduce interactive interfaces in agent design to enhance their adaptability in real-world applications.
Conclusion
The release of XiangqiBench provides new insights and methods for evaluating LLM agents, emphasizing the importance of closed-loop evaluation and offering valuable experimental data for AI researchers and developers.
— END —Source: ArXiv NLP/LLM (cs.CL) (2026-10-05)
Tags: #XiangqiBench #LLM Evaluation #Closed-loop Agents #arXiv
Community Comments