English
Related papers

Related papers: BSODiag: A Global Diagnosis Framework for Batch Se…

200 papers

Integrated cyber-physical systems (CPSs), such as the smart grid, are increasingly becoming the underpinning technology for major industries. A major concern regarding such systems are the seemingly unexpected large-scale failures, which…

Physics and Society · Physics 2018-07-27 Yingrui Zhang , Osman Yagan

Performance unpredictability in cloud services leads to poor user experience, degraded availability, and has revenue ramifications. Detecting performance degradation a posteriori helps the system take corrective action, but does not avoid…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-04-25 Yu Gan , Meghna Pancholi , Dailun Cheng , Siyuan Hu , Yuan He , Christina Delimitrou

Modern large-scale data-farms consist of hundreds of thousands of storage devices that span distributed infrastructure. Devices used in modern data centers (such as controllers, links, SSD- and HDD-disks) can fail due to hardware as well as…

Managing access between large numbers of distributed medical devices has become a crucial aspect of modern healthcare systems, enabling the establishment of smart hospitals and telehealth infrastructure. However, as telehealth technology…

Cryptography and Security · Computer Science 2023-12-07 Khalid Al-hammuri , Fayez Gebali , Awos Kanan

Cloud systems are susceptible to performance issues, which may cause service-level agreement violations and financial losses. In current practice, crucial metrics are monitored periodically to provide insight into the operational status of…

Machine Learning · Computer Science 2024-11-08 Wenwei Gu , Jinyang Liu , Zhuangbin Chen , Jianping Zhang , Yuxin Su , Jiazhen Gu , Cong Feng , Zengyin Yang , Yongqiang Yang , Michael Lyu

In today's global economy, supply chain (SC) entities have become increasingly interconnected with demand and supply relationships due to the need for strategic outsourcing. Such interdependence among firms not only increases efficiency but…

Physics and Society · Physics 2020-11-16 Qihui Yang , Caterina Scoglio , Don Gruenbacher

AI workloads incur frequent failures and incidents from the underlying infrastructure. The current incident management workflow follows a provider-centric paradigm, where users report incidents to the infrastructure provider who then…

Software Engineering · Computer Science 2026-05-08 Yitao Yang , Yangtao Deng , Yifan Xiong , Baochun Li , Hong Xu , Peng Cheng

While various service orchestration aspects within Computing Continuum (CC) systems have been extensively addressed, including service placement, replication, and scheduling, an open challenge lies in ensuring uninterrupted data delivery…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-02 Ivan Čilić , Valentin Jukanović , Ivana Podnar Žarko , Pantelis Frangoudis , Schahram Dustdar

Cloud services have recently started undergoing a major shift from monolithic applications, to graphs of hundreds of loosely-coupled microservices. Microservices fundamentally change a lot of assumptions current cloud systems are designed…

Kubernetes has emerged as an essential platform for deploying containerised applications across cloud and edge infrastructures. As Kubernetes gains increasing adoption for mission-critical microservices, evaluating system resilience under…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-07-23 Zihao Chen , Mohammad Goudarzi , Adel Nadjaran Toosi

The presence of unhealthy nodes in cloud infrastructure signals the potential failure of machines, which can significantly impact the availability and reliability of cloud services, resulting in negative customer experiences. Effectively…

Systems and Control · Electrical Eng. & Systems 2024-10-24 Chaoyun Zhang , Randolph Yao , Si Qin , Ze Li , Shekhar Agrawal , Binit R. Mishra , Tri Tran , Minghua Ma , Qingwei Lin , Murali Chintalapati , Dongmei Zhang

This paper introduces a scalable Anomaly Detection Service with a generalizable API tailored for industrial time-series data, designed to assist Site Reliability Engineers (SREs) in managing cloud infrastructure. The service enables…

Machine Learning · Computer Science 2025-01-29 Nimesh Jha , Shuxin Lin , Srideepika Jayaraman , Kyle Frohling , Christodoulos Constantinides , Dhaval Patel

Root cause analysis in modern cloud infrastructure demands sophisticated understanding of heterogeneous data sources, particularly time-series performance metrics that involve core failure signatures. While large language models demonstrate…

Artificial Intelligence · Computer Science 2026-01-09 Gijun Park

As business of Alibaba expands across the world among various industries, higher standards are imposed on the service quality and reliability of big data cloud computing platforms which constitute the infrastructure of Alibaba Cloud.…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-11-09 Yingying Zhang , Zhengxiong Guan , Huajie Qian , Leili Xu , Hengbo Liu , Qingsong Wen , Liang Sun , Junwei Jiang , Lunting Fan , Min Ke

Performance unpredictability is a major roadblock towards cloud adoption, and has performance, cost, and revenue ramifications. Predictable performance is even more critical as cloud services transition from monolithic designs to…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-05-06 Yu Gan , Yanqi Zhang , Kelvin Hu , Dailun Cheng , Yuan He , Meghna Pancholi , Christina Delimitrou

Many large-scale software systems demonstrate metastable failures. In this class of failures, a stressor such as a temporary spike in workload causes the system performance to drop and, subsequently, the system performance continues to…

Extreme Edge Computing (XEC) distributes streaming workloads across consumer-owned devices, exploiting their proximity to users and ubiquitous availability. Many such workloads are AI-driven, requiring continuous neural network inference…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-19 MHD Saria Allahham , Hossam S. Hassanein

With the growing complexity of cyberattacks targeting critical infrastructures such as water treatment networks, there is a pressing need for robust anomaly detection strategies that account for both system vulnerabilities and evolving…

Machine Learning · Computer Science 2025-08-14 Arun Vignesh Malarkkan , Haoyue Bai , Dongjie Wang , Yanjie Fu

Caching at the edge of wireless networks is a keytechnology to reduce traffic in the backhaul link. However, aconcentrated amount of requests during peak-periods may causethe outage of the system, meaning that the network is not ableto…

Networking and Internet Architecture · Computer Science 2021-01-05 Estefanía Recayte , Andrea Munari

Detecting failures and identifying their root causes promptly and accurately is crucial for ensuring the availability of microservice systems. A typical failure troubleshooting pipeline for microservices consists of two phases: anomaly…

Software Engineering · Computer Science 2024-05-16 Luan Pham , Huong Ha , Hongyu Zhang