English

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Artificial Intelligence 2026-07-01 v1

Abstract

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception, Interaction, Reasoning, and Internalization, and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning(SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github. com/LAMDA-NeSy/OpenAgent.

Keywords

Cite

@article{arxiv.2607.01084,
  title  = {Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use},
  author = {Song-Lin Lv and Weiming Wu and Rui Zhu and Zi-Jian Cheng and Lan-Zhe Guo},
  journal= {arXiv preprint arXiv:2607.01084},
  year   = {2026}
}

Comments

Accepted by ICML 2026