Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer from being largely episodic, domain-specific, or insufficiently grounded in persistent economic dynamics. We introduce EcoGym, a generalizable benchmark for continuous plan-and-execute decision making in interactive economies. EcoGym comprises three diverse environments: Vending (adapted from the closed-source Vending-Bench, with full open-source release), Freelance (new), and Operation (new), implemented in a unified decision-making process with standardized interfaces, and budgeted actions over an effectively unbounded horizon (1000+ steps if 365 day-loops for evaluation). The evaluation of EcoGym is based on business-relevant outcomes (e.g., net worth, income, and DAU), targeting long-term strategic coherence and robustness under partial observability and stochasticity. Experiments across eleven leading LLMs expose a systematic tension: no single model dominates across all three scenarios. Critically, we find that models exhibit significant suboptimality in either high-level strategies or efficient actions executions. EcoGym is released as an open, extensible testbed for transparent long-horizon agent evaluation and for studying controllability utility trade-offs in economic settings.
@article{arxiv.2602.09514,
title = {EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies},
author = {Xavier Hu and Jinxiang Xia and Shengze Xu and Kangqi Song and Yishuo Yuan and Guibin Zhang and JinCheng Ren and Boyu Feng and Li Lu and Tieyong Zeng and Jiaheng Liu and Minghao Liu and He Zhu and Yuchen Eleanor Jiang and Wei Wang and Wangchunshu Zhou},
journal= {arXiv preprint arXiv:2602.09514},
year = {2026}
}