English

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Computation and Language 2026-05-26 v2 Artificial Intelligence

Abstract

Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Model Context Protocol (MCP) offers a unified interface to discover and invoke tools dynamically. However, there is a significant gap in benchmarking multi-step tasks using diverse MCP tools in realistic, dynamic scenarios. In this work, we present LiveMCP-101, a benchmark of 101 real-world queries that require coordinated use of multiple MCP tools. To address temporal variability in real-world tool responses, we introduce a parallel evaluation framework where a reference agent executes a validated plan simultaneously to produce real-time reference outputs. Experiments show that even frontier LLMs achieve a success rate below 60\%, highlighting challenges in multi-step tool use. Comprehensive error analysis identifies seven failure modes spanning tool planning, parameterization, and output handling, pointing to concrete directions for improving current models. LiveMCP-101 sets a rigorous standard for evaluating real-world agent capabilities, advancing toward autonomous agent systems that reliably execute complex tasks through MCP tool orchestration.

Keywords

Cite

@article{arxiv.2508.15760,
  title  = {LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries},
  author = {Ming Yin and Dinghan Shen and Silei Xu and Sixun Dong and Mian Zhang and Yebowen Hu and Shujian Liu and Jianbing Han and Simin Ma and Song Wang and Sathish Reddy Indurthi and Xun Wang and Yiran Chen and Kaiqiang Song},
  journal= {arXiv preprint arXiv:2508.15760},
  year   = {2026}
}
R2 v1 2026-07-01T05:00:32.356Z