English

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V

Artificial Intelligence 2026-04-10 v1

Abstract

Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation signals rich enough for long-horizon, multi-agent play. We introduce CivBench, a benchmark for LLM strategists (i.e., agentic setups) in multiplayer Civilization V. Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout play, validated through predictive, construct, and convergent validity. Across 307 games with 7 LLMs and multiple CivBench agent conditions, we demonstrate CivBench's potential to estimate strategic capabilities as an unsaturated benchmark, reveal model-specific effects of agentic setup, and outline distinct strategic profiles not visible through outcome-only evaluation.

Keywords

Cite

@article{arxiv.2604.07733,
  title  = {CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V},
  author = {John Chen and Sihan Cheng and Can Gurkan and Mingyi Lin},
  journal= {arXiv preprint arXiv:2604.07733},
  year   = {2026}
}

Comments

Under review

R2 v1 2026-07-01T12:00:25.519Z