English

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

Artificial Intelligence 2026-05-19 v2

Abstract

Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks with immediate feedback, this setting requires agents to make binding commitments under partial observability, delayed consequences, hard resource budgets, and shifting dynamics. We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes. The simulator is built from transformed firm-level financial data, anonymized business documents, decade-scale macroeconomic and industry signals, and expert-validated operating rules. Experiments across 23 LLMs and four agent frameworks show that current agents remain far from robust: only 15.4% of trials survive the full horizon, larger models do not reliably outperform smaller ones, and failures cascade across observation, action timing, and capital sizing. These findings establish long-horizon resource allocation under uncertainty as a distinct capability gap for LLM agents.

Keywords

Cite

@article{arxiv.2603.23638,
  title  = {Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment},
  author = {Yi Han and Yan Wang and Lingfei Qian and Haohang Li and Yupeng Cao and Yueru He and Xueqing Peng and Nanhan Shen and Yitao Xu and Yankai Chen and Dongji Feng and Jimin Huang and Xue Liu and Jian-Yun Nie and Sophia Ananiadou},
  journal= {arXiv preprint arXiv:2603.23638},
  year   = {2026}
}