English

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

Artificial Intelligence 2026-07-06 v1

Abstract

Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They typically rely on pre-packaged tool interfaces, overlook critical steps, and assume inputs are clean and fully specified. Consequently, they understate the difficulty of real deployments, where uncertainty and noise are ubiquitous and agents must proactively explore the environment to uncover new tools. To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands. Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information. Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2, revealing a substantial gap between the capability of current agents and the demands of real-world applications.

Keywords

Cite

@article{arxiv.2607.05174,
  title  = {AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments},
  author = {Zhiheng Xi and Dingwen Yang and Jiaqi Liu and Jixuan Huang and Honglin Guo and Baodai Huang and Tinggang Chen and Qi Zhang and Zhonghang Lu and Chenyu Liu and Jiajun Sun and Jiazheng Zhang and Dingwei Zhu and Xin Guo and Junzhe Wang and Zhihao Zhang and Yuming Yang and Junjie Ye and Minghe Gao and Dongrui Liu and Jiaming Ji and Guohao Li and Tao Gui and Qi Zhang and Xuanjing Huang},
  journal= {arXiv preprint arXiv:2607.05174},
  year   = {2026}
}

Comments

Accepted as a main conference paper at ACL 2026