ZICQ
中 Log in / Sign up
ZICQ Info LLMs & Foundation Models #Vox Deorum #CivBench #LLMs & Foundation Models #Strategy Game #AI Benchmark

Vox Deorum Releases CivBench: Benchmarking LLMs in Playing Civilization V

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Vox Deorum has introduced CivBench, a benchmark designed to evaluate the performance of large language models (LLMs) in the turn-based strategy game Civilization V. This controlled experiment rotates LLMs through three fixed starts, testing their capabilities in long-term decision-making and complex strategic tasks. Initial results show that models like GLM-5.3, Opus-5.5, and Qwen-3.8-27B perform well, with GLM-5.3 achieving a cultural victory and Qwen-3.8-27B excelling in a science victory. Civ


CivBench: Evaluating LLM Strategy in Civilization V

Vox Deorum has introduced CivBench, a benchmark tool designed to evaluate the performance of large language models (LLMs) in the turn-based strategy game Civilization V. The tool offers the following features to assess LLM strategic capabilities:

  • Controlled Experiment Environment: CivBench uses three fixed starting conditions to ensure that each model is tested under the same circumstances, enhancing result comparability.
  • Multi-Civilization Participation: Each game includes eight civilizations, with two controlled by the tested LLM and the remaining six by the standard Vox Populi AI. The LLM is responsible for high-level strategy, while the game's built-in AI handles low-level execution.
  • Diverse Victory Conditions: The tool tests models on different victory conditions, including cultural and scientific victories.

Key Features

  • Reproducibility: By using fixed starting conditions and multi-civilization participation, CivBench ensures the reproducibility of test results.
  • Openness: Vox Deorum is open-source, allowing users to download and use the tool for their own tests.
  • Multi-Model Support: It supports various LLMs, including local OpenAI-compatible servers and Qwen-3.8-27B.

Test Results

Preliminary results show that GLM-5.3 excels in achieving a cultural victory, while Qwen-3.8-27B demonstrates strong performance in achieving a scientific victory. Here are some specific test cases:

  • GLM-5.3: Achieved a cultural victory as the Chinese civilization.
  • Opus-5.5: Achieved a cultural victory as the Moroccan civilization.
  • Qwen-3.8-27B: Achieved a scientific victory as the Chinese civilization.

Developer Recommendations

  • Participate in Testing: Developers can download Vox Deorum and use their own LLMs for testing, or participate in online games with Vox Deorum to compete or collaborate with LLMs.
  • Feedback and Improvement: The Vox Deorum team welcomes user feedback and plans to add more features and model support in future versions.

Industry Impact

The release of CivBench provides researchers and developers with a new platform to evaluate and compare LLM performance in complex tasks. This not only helps to understand LLM performance in strategy games but also offers new insights into the application of AI in simulating complex decision-making scenarios in the real world.

Conclusion

The introduction of CivBench marks a further development in AI benchmarking tools, providing a new evaluation standard for LLM performance in strategy games. As more models are added and testing deepens, CivBench is expected to become an important tool in the AI research field.


Source: Reddit r/LocalLLaMA (2026-10-05)

— END —

Tags: #Vox Deorum #CivBench #LLMs & Foundation Models #Strategy Game #AI Benchmark

Community Comments

Loading live comments and annotations…