English

GuardianAgentBench: Where Agents Fail and How to Guard Them

Artificial Intelligence 2026-07-23 v1

Abstract

As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three production-ready frameworks: LangChain, LlamaIndex, and Vectara. The benchmark incorporates rigorous multi-stage validation and five adversarial attack modes. Experiments with six state-of-the-art models reveal that even the strongest configuration achieves only 74.8% overall accuracy and expose two distinct failure regimes: stronger models under-call required tools, while weaker models mis-select and over-call tools. Performance degrades monotonically with both tool-set size and sequential turn depth, with long-horizon planning proving the steeper bottleneck. Our guardrail implementation consistently outperforms system-prompt-based defenses across all models, recovering 19.9% of failures at a false positive rate of just 0.5%. These results demonstrate that execution-time structural intervention improves safety without disrupting correct agent behavior.

Keywords

Cite

@article{arxiv.2607.20982,
  title  = {GuardianAgentBench: Where Agents Fail and How to Guard Them},
  author = {Vishal Ishwar Naik and Chenyu Xu and Donna Dong and Hussein Hassan and Abhishek Pradhan and Ofer Mendelevitch and Tallat Shafat and Humayun Irshad},
  journal= {arXiv preprint arXiv:2607.20982},
  year   = {2026}
}