English

EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

Computation and Language 2025-05-27 v2 Artificial Intelligence Machine Learning

Abstract

Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative adaptation in unfamiliar environments. To address this, we introduce EscapeBench, a benchmark suite of room escape game environments designed to challenge agents with creative reasoning, unconventional tool use, and iterative problem-solving to uncover implicit goals. Our results show that current LM models, despite employing working memory and Chain-of-Thought reasoning, achieve only 15% average progress without hints, highlighting their limitations in creativity. To bridge this gap, we propose EscapeAgent, a framework designed to enhance creative reasoning through Foresight (innovative tool use) and Reflection (identifying unsolved tasks). Experiments show that EscapeAgent can execute action chains over 1,000 steps while maintaining logical coherence. It navigates and completes games with up to 40% fewer steps and hints, performs robustly across difficulty levels, and achieves higher action success rates with more efficient and innovative puzzle-solving strategies.

Keywords

Cite

@article{arxiv.2412.13549,
  title  = {EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents},
  author = {Cheng Qian and Peixuan Han and Qinyu Luo and Bingxiang He and Xiusi Chen and Yuji Zhang and Hongyi Du and Jiarui Yao and Xiaocheng Yang and Denghui Zhang and Yunzhu Li and Heng Ji},
  journal= {arXiv preprint arXiv:2412.13549},
  year   = {2025}
}

Comments

23 pages, 15 figures, ACL 2025 Main Conference

R2 v1 2026-06-28T20:39:57.112Z