English
Related papers

Related papers: Assess and Summarize: Improve Outage Understanding…

200 papers

Incident management for cloud services is a complex process involving several steps and has a huge impact on both service health and developer productivity. On-call engineers require significant amount of domain knowledge and manual effort…

Software Engineering · Computer Science 2023-02-10 Toufique Ahmed , Supriyo Ghosh , Chetan Bansal , Thomas Zimmermann , Xuchao Zhang , Saravan Rajmohan

Cloud services are omnipresent and critical cloud service failure is a fact of life. In order to retain customers and prevent revenue loss, it is important to provide high reliability guarantees for these services. One way to do this is by…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-11-13 Shubham Agarwal , Sarthak Chakraborty , Shaddy Garg , Sumit Bisht , Chahat Jain , Ashritha Gonuguntla , Shiv Saini

Incident management is a key aspect of operating large-scale cloud services. To aid with faster and efficient resolution of incidents, engineering teams document frequent troubleshooting steps in the form of Troubleshooting Guides (TSGs),…

Software Engineering · Computer Science 2022-05-27 Manish Shetty , Chetan Bansal , Sai Pramod Upadhyayula , Arjun Radhakrishna , Anurag Gupta

Inadequate service availability is the top concern when employing Cloud computing. It has been recognized that zero downtime is impossible for large-scale Internet services. By learning from the previous and others' mistakes, nevertheless,…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-12-24 Zheng Li , Mingfei Liang , Liam O'Brien , He Zhang

Despite significant reliability efforts, large-scale cloud services inevitably experience production incidents that can significantly impact service availability and customer's satisfaction. Worse, in many cases one incident can lead to…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-03-28 Supriyo Ghosh , Karish Grover , Jimmy Wong , Chetan Bansal , Rakesh Namineni , Mohit Verma , Saravan Rajmohan

Modern cloud services are prone to failures due to their complex architecture, making diagnosis a critical process. Site Reliability Engineers (SREs) spend hours leveraging multiple sources of data, including the alerts, error logs, and…

Software Engineering · Computer Science 2023-09-15 Sarthak Chakraborty , Shubham Agarwal , Shaddy Garg , Abhimanyu Sethia , Udit Narayan Pandey , Videh Aggarwal , Shiv Saini

Cloud-based services are surging into popularity in recent years. However, outages, i.e., severe incidents that always impact multiple services, can dramatically affect user experience and incur severe economic losses. Locating the…

Incident management for large cloud services is a complex and tedious process and requires significant amount of manual efforts from on-call engineers (OCEs). OCEs typically leverage data from different stages of the software development…

Networking and Internet Architecture · Computer Science 2024-04-08 Drishti Goel , Fiza Husain , Aditya Singh , Supriyo Ghosh , Anjaly Parayil , Chetan Bansal , Xuchao Zhang , Saravan Rajmohan

In order to plan for failure recovery, the designers of cloud systems need to understand how their system can potentially fail. Unfortunately, analyzing the failure behavior of such systems can be very difficult and time-consuming, due to…

Software Engineering · Computer Science 2022-03-09 Domenico Cotroneo , Luigi De Simone , Pietro Liguori , Roberto Natella , Nematollah Bidokhti

Cloud computing has fundamentally transformed application development, yet a gap remains between the serverless promise of simplified deployment and its practical realization due to fragmentation across function runtimes, state management,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-30 Pawissanutt Lertpongrujikorn

For Open Source Software (OSS) projects, discussions in Issue Tracking Systems (ITS) serve as a crucial collaboration mechanism for diverse stakeholders. However, these discussions can become lengthy and entangled, making it hard to find…

Human-Computer Interaction · Computer Science 2023-08-08 Saskia Gilmer , Avinash Bhat , Shuvam Shah , Kevin Cherry , Jinghui Cheng , Jin L. C. Guo

In recent decades, the weather around the world has become more irregular and extreme, often causing large-scale extended power outages. Resilience -- the capability of withstanding, adapting to, and recovering from a large-scale disruption…

Applications · Statistics 2025-08-06 Shixiang Zhu , Rui Yao , Yao Xie , Feng Qiu , Yueming Qiu , Xuan Wu

The presence of unhealthy nodes in cloud infrastructure signals the potential failure of machines, which can significantly impact the availability and reliability of cloud services, resulting in negative customer experiences. Effectively…

Systems and Control · Electrical Eng. & Systems 2024-10-24 Chaoyun Zhang , Randolph Yao , Si Qin , Ze Li , Shekhar Agrawal , Binit R. Mishra , Tri Tran , Minghua Ma , Qingwei Lin , Murali Chintalapati , Dongmei Zhang

Runtime failure and performance degradation is commonplace in modern cloud systems. For cloud providers, automatically determining the root cause of incidents is paramount to ensuring high reliability and availability as prompt fault…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-12 Zhiqiang Xie , Yujia Zheng , Lizi Ottens , Kun Zhang , Christos Kozyrakis , Jonathan Mace

Social media analysis of disaster events is a critical task in crisis informatics research. It involves analyzing social media data generated during natural disasters, crisis events, or other mass convergence events. Due to the large data…

Software Engineering · Computer Science 2020-07-09 Gerard Casas Saez

This paper presents the results of a research study related to software system failures, with the goal of understanding how we might better evolve, maintain and support software systems in production. We have qualitatively analyzed thirty…

Software Engineering · Computer Science 2020-08-26 Jonathan Sillito , Esdras Kutomi

Reviews are central to how travelers evaluate products on online marketplaces, yet existing summarization research often emphasizes end-to-end quality while overlooking benchmark reliability and the practical utility of granular insights.…

Computation and Language · Computer Science 2026-03-23 Piyush Kumar Singh , Jayesh Choudhari

Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis-small, critical, and easily lost in a desert…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Zhijia Liang , Jiaming Li , Weikai Chen , Yanhao Zhang , Haonan Lu , Guanbin Li

Cloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages…

Software Engineering · Computer Science 2019-09-04 Domenico Cotroneo , Luigi De Simone , Pietro Liguori , Roberto Natella , Nematollah Bidokhti

Cloud infrastructure is the collective term for all physical devices within cloud systems. Failures within the cloud infrastructure system can severely compromise the stability and availability of cloud services. Particularly, batch servers…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-02-25 Tao Duan , Runqing Chen , Pinghui Wang , Junzhou Zhao , Jiongzhou Liu , Shujie Han , Yi Liu , Fan Xu
‹ Prev 1 2 3 10 Next ›