English
Related papers

Related papers: Praxium: Diagnosing Cloud Anomalies with AI-based …

200 papers

Root Cause Analysis (RCA) plays a pivotal role in the incident diagnosis process for cloud services, requiring on-call engineers to identify the primary issues and implement corrective actions to prevent future recurrences. Improving the…

Computation and Language · Computer Science 2024-01-26 Xuchao Zhang , Supriyo Ghosh , Chetan Bansal , Rujia Wang , Minghua Ma , Yu Kang , Saravan Rajmohan

Root cause analysis (RCA) is essential for diagnosing failures within complex software systems to ensure system reliability. The highly distributed and interdependent nature of modern cloud-based systems often complicates RCA efforts,…

Software Engineering · Computer Science 2026-02-02 Evelien Riddell , James Riddell , Gengyi Sun , Michał Antkiewicz , Krzysztof Czarnecki

While microservices are revolutionizing cloud computing by offering unparalleled scalability and independent deployment, their decentralized nature poses significant security and management challenges that can threaten system stability. We…

Software Engineering · Computer Science 2025-06-30 Matteo Esposito , Alexander Bakhtin , Noman Ahmad , Mikel Robredo , Ruoyu Su , Valentina Lenarduzzi , Davide Taibi

In the modern world, we are permanently using, leveraging, interacting with, and relying upon systems of ever higher sophistication, ranging from our cars, recommender systems in e-commerce, and networks when we go online, to integrated…

Artificial Intelligence · Computer Science 2023-06-23 Patrick Rodler

The persistently growing resilience concerns of large-scale computing systems today require not only generic fault tolerance approaches, but also application-level resilience, due to demanding efficiency and various domain-specific…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-11-07 Li Tan , Marc Charest , Nathan DeBardeleben , Qiang Guan , Ben Bergen

As more and more companies are migrating (or planning to migrate) from on-premise to Cloud, their focus is to find anomalies and deficits as early as possible in the development life cycle. We propose Frisbee, a declarative language and…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-09-23 Fotis Nikolaidis , Antony Chazapis , Manolis Marazakis , Angelos Bilas

Cloud-enabled large-scale distributed systems orchestrate resources and services from various providers in order to deliver high-quality software solutions to the end users. The space and structure created by such technological advancements…

Software Engineering · Computer Science 2018-08-14 Andreea Buga , Sorana Tania Nemes , Atif Mashkoor

Anomaly detection in cloud environments remains both critical and challenging. Existing context-level benchmarks typically focus on either metrics or logs and often lack reliable annotation, while most detection methods emphasize point…

Artificial Intelligence · Computer Science 2025-10-07 Xinkai Zou , Xuan Jiang , Ruikai Huang , Haoze He , Parv Kapoor , Hongrui Wu , Yibo Wang , Jian Sha , Xiongbo Shi , Zixun Huang , Jinhua Zhao

Designing software compatible with cloud-based Microservice Architectures (MSAs) is vital due to the performance, scalability, and availability limitations. As the complexity of a system increases, it is subject to deprecation, difficulties…

Software Engineering · Computer Science 2024-07-22 Thakshila Imiya Mohottige , Artem Polyvyanyy , Rajkumar Buyya , Colin Fidge , Alistair Barros

Microservice based systems underpin modern distributed computing environments but remain vulnerable to partial failures, cascading timeouts, and inconsistent recovery behavior. Although numerous resilience and recovery patterns have been…

Software Engineering · Computer Science 2026-02-03 Muzeeb Mohammad

Anomaly detection in complex, high-dimensional data, such as UAV sensor readings, is essential for operational safety but challenging for existing methods due to their limited sensitivity, scalability, and inability to capture intricate…

Machine Learning · Computer Science 2025-10-28 Mingze Gong , Juan Du , Jianbang You

Modern machine learning workloads such as large language model training, fine-tuning jobs are highly distributed and span across hundreds of systems with multiple GPUs. Job completion time for these workloads is the artifact of the…

Networking and Internet Architecture · Computer Science 2025-07-08 Jit Gupta , Tarun Banka , Rahul Gupta , Mithun Dharmaraj , Jasleen Kaur

Serverless applications can be particularly difficult to troubleshoot, as these applications are often composed of various managed and partly managed services. Faults are often unpredictable and can occur at multiple points, even in simple…

Software Engineering · Computer Science 2024-07-16 Maria C. Borges , Sebastian Werner , Ahmet Kilic

Agentic AI, with goal-directed, proactive, and autonomous decision-making capabilities, offers a compelling opportunity to address movement-related risks in human activity, including the persistent hazard of falls among elderly populations.…

Artificial Intelligence · Computer Science 2026-04-22 Farbod Zorriassatine , Ahmad Lotfi

Large Language Model (LLM)-based agents increasingly rely on APIs to operate complex web applications, but rapid evolution often leads to incomplete or inconsistent API documentation. Existing work falls into two categories: (1) static,…

Software Engineering · Computer Science 2026-03-26 Yanjing Yang , Chenxing Zhong , Ke Han , Zeru Cheng , Jinwei Xu , Xin Zhou , He Zhang , Bohan Liu

While cloud environments and auto-scaling solutions have been widely applied to traditional monolithic applications, they face significant limitations when it comes to microservices-based architectures. Microservices introduce additional…

Software Engineering · Computer Science 2025-02-03 Majid Dashtbani , Ladan Tahvildari

Anomaly detection significantly enhances the robustness of cloud systems. While neural network-based methods have recently demonstrated strong advantages, they encounter practical challenges in cloud environments: the contradiction between…

Machine Learning · Computer Science 2024-03-20 Feiyi Chen , Yingying zhang , Zhen Qin , Lunting Fan , Renhe Jiang , Yuxuan Liang , Qingsong Wen , Shuiguang Deng

Reliability is extremely important for large-scale cloud systems like Microsoft 365. Cloud failures such as disk failure, node failure, etc. threaten service reliability, resulting in online service interruptions and economic loss. Existing…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-09-07 Fangkai Yang , Wenjie Yin , Lu Wang , Tianci Li , Pu Zhao , Bo Liu , Paul Wang , Bo Qiao , Yudong Liu , Mårten Björkman , Saravan Rajmohan , Qingwei Lin , Dongmei Zhang

High-Performance Computing (HPC) centers and cloud providers support an increasingly diverse set of applications on heterogenous hardware. As Artificial Intelligence (AI) and Machine Learning (ML) workloads have become an increasingly…

Ensuring the reliability and availability of cloud services necessitates efficient root cause analysis (RCA) for cloud incidents. Traditional RCA methods, which rely on manual investigations of data sources such as logs and traces, are…

‹ Prev 1 8 9 10 Next ›