English

CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments

Machine Learning 2026-03-31 v1 Artificial Intelligence Information Retrieval Performance

Abstract

The increasing agentic capabilities of Large Language Models (LLMs) have enabled their deployment in real-world applications, such as cloud services, where customer-assistant interactions exhibit high technical complexity and long-horizon dependencies, making robustness and resolution efficiency critical for customer satisfaction. However, existing benchmarks for LLM-based agents largely rely on synthetic environments that fail to capture the diversity and unpredictability of authentic customer inputs, often ignoring the resolution efficiency essential for real-world deployment. To bridge this gap, we introduce CirrusBench, a novel evaluation framework distinguished by its foundation in real-world data from authentic cloud service tickets. CirrusBench preserves the intricate multi-turn logical chains and realistic tool dependencies inherent to technical service environments. Moving beyond execution correctness, we introduce novel Customer-Centric metrics to define agent success, quantifying service quality through metrics such as the Normalized Efficiency Index and Multi-Turn Latency to explicitly measure resolution efficiency. Experiments utilizing our framework reveal that while state-of-the-art models demonstrate strong reasoning capabilities, they frequently struggle in complex, realistic multi-turn tasks and fail to meet the high-efficiency standards required for customer service, highlighting critical directions for the future development of LLM-based agents in practical technical service applications. CirrusBench evaluation framework is released at: https://github.com/CirrusAI

Keywords

Cite

@article{arxiv.2603.28569,
  title  = {CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments},
  author = {Yi Yu and Guangquan Hu and Chenghuang Shen and Xingyan Liu and Jing Gu and Hangyi Sun and Junzhuo Ma and Weiting Liu and Jianfeng Liu and Mingyue Pu and Yu Wang and Zhengdong Xiao and Rui Xie and Longjiu Luo and Qianrong Wang and Gurong Cui and Honglin Qiao and Wenlian Lu},
  journal= {arXiv preprint arXiv:2603.28569},
  year   = {2026}
}

Comments

Submitted for SIGKDD 2026

R2 v1 2026-07-01T11:44:19.048Z