English
Related papers

Related papers: BSODiag: A Global Diagnosis Framework for Batch Se…

200 papers

Cloud computing provides resources over the Internet and allows a plethora of applications to be deployed to provide services for different industries. The major bottleneck being faced currently in these cloud frameworks is their limited…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-11-18 Shreshth Tuli , Nipam Basumatary , Sukhpal Singh Gill , Mohsen Kahani , Rajesh Chand Arya , Gurpreet Singh Wander , Rajkumar Buyya

DRAM failure prediction is a vital task in AIOps, which is crucial to maintain the reliability and sustainable service of large-scale data centers. However, limited work has been done on DRAM failure prediction mainly due to the lack of…

Machine Learning · Computer Science 2021-05-05 Zhiyue Wu , Hongzuo Xu , Guansong Pang , Fengyuan Yu , Yijie Wang , Songlei Jian , Yongjun Wang

Serverless computing has redefined cloud application deployment by abstracting infrastructure and enabling on-demand, event-driven execution, thereby enhancing developer agility and scalability. However, maintaining consistent application…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-01 Chanh Nguyen , Erik Elmroth , Monowar Bhuyan

Cascading failures (CF) entail component breakdowns spreading through infrastructure networks, causing system-wide collapse. Predicting CFs is of great importance for infrastructure stability and urban function. Despite extensive research…

Social and Information Networks · Computer Science 2025-03-06 Yinzhou Tang , Jinghua Piao , Huandong Wang , Shaw Rajib , Yong Li

Telecommunications networks rely on configurations to define routing behavior, especially in the Border Gateway Protocol (BGP), where misconfigurations can lead to severe outages and security breaches, as demonstrated by the 2021 Facebook…

Software Engineering · Computer Science 2025-12-08 Chenlu Zhang , Amirmohammad Pasdar , Van-Thuan Pham

Services hosted in multi-tenant cloud platforms often encounter performance interference due to contention for non-partitionable resources, which in turn causes unpredictable behavior and degradation in application performance. To grapple…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-04-15 Yogesh D. Barve , Shashank Shekhar , Ajay Dev Chhokra , Shweta Khare , Anirban Bhattacharjee , Zhuangwei Kang , Hongyang Sun , Aniruddha Gokhale

Operating a modern power grid reliably in case of SCADA/EMS failure or amid difficult times like COVID-19 pandemic is a challenging task for grid operators. In [11], a PMU-based emergency generation dispatch scheme has been proposed to help…

Physics and Society · Physics 2021-08-03 Song Zhang , Xiaochuan Luo , Eugene Litvinov

Automatic failure diagnosis is crucial for large microservice systems. Currently, most failure diagnosis methods rely solely on single-modal data (i.e., using either metrics, logs, or traces). In this study, we conduct an empirical study…

Software Engineering · Computer Science 2023-06-01 Shenglin Zhang , Pengxiang Jin , Zihan Lin , Yongqian Sun , Bicheng Zhang , Sibo Xia , Zhengdan Li , Zhenyu Zhong , Minghua Ma , Wa Jin , Dai Zhang , Zhenyu Zhu , Dan Pei

Low latency and high availability of an app or a web service are key, amongst other factors, to the overall user experience (which in turn directly impacts the bottomline). Exogenic and/or endogenic factors often give rise to breakouts in…

Methodology · Statistics 2014-12-01 Nicholas A. James , Arun Kejariwal , David S. Matteson

Cloud computing recently developed into a viable alternative to on-premises systems for executing high-performance computing (HPC) applications. With the emergence of new vendors and hardware options, there is now a growing need to…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-12-14 Mohammad Mohammadi , Timur Bazhirov

Due to the growing complexity of modern data centers, failures are not uncommon any more. Therefore, fault tolerance mechanisms play a vital role in fulfilling the availability requirements. Multiple availability models have been proposed…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-26 Otto Bibartiu , Frank Dürr , Kurt Rothermel , Beate Ottenwälder , Andreas Grau

Simulating potential cascading failures can be useful for avoiding or mitigating such events. Currently, existing steady-state analysis tools are ill-suited for simulating cascading outages as they do not model frequency dependencies, they…

Signal Processing · Electrical Eng. & Systems 2019-11-25 Amritanshu Pandey , Aayushya Agarwal , Marko Jereminov , Martin R. Wagner , David M. Bromberg , Larry Pileggi

The exponential growth of Internet of Things (IoT) has given rise to a new wave of edge computing due to the need to process data on the edge, closer to where it is being produced and attempting to move away from a cloud-centric…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-11-15 Hamza Javed , Adel N. Toosi , Mohammad S. Aslanpour

In the data center, unexpected downtime caused by memory failures can lead to a decline in the stability of the server and even the entire information technology infrastructure, which harms the business. Therefore, whether the memory…

Databases · Computer Science 2021-05-18 Chengdong Yao

Cloud applications are increasingly shifting from large monolithic services to complex graphs of loosely-coupled microservices. Despite the advantages of modularity and elasticity microservices offer, they also complicate cluster management…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-01-05 Yu Gan , Mingyu Liang , Sundar Dev , David Lo , Christina Delimitrou

Cloud computing is ubiquitous: more and more companies are moving the workloads into the Cloud. However, this rise in popularity challenges Cloud service providers, as they need to monitor the quality of their ever-growing offerings…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-08-04 Mohammad Saiful Islam , William Pourmajidi , Lei Zhang , John Steinbacher , Tony Erwin , Andriy Miranskyy

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam…

Machine Learning · Computer Science 2025-01-06 Milan Jain , Burcu O. Mutlu , Caleb Stam , Jan Strube , Brian A. Schupbach , Jason M. St. John , William A. Pellico

Serverless computing has become a major trend among cloud providers. With serverless computing, developers fully delegate the task of managing the servers, dynamically allocating the required resources, as well as handling availability and…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-06-08 Pascal Maissen , Pascal Felber , Peter Kropf , Valerio Schiavoni

Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We populate this repository with the data we collect containing…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-15 Sacheendra Talluri , Dante Niewenhuis , Xiaoyu Chu , Jakob Kyselica , Mehmet Cetin , Alexander Balgavy , Alexandru Iosup

Monitoring is an essential aspect of maintaining and developing computer systems that increases in difficulty proportional to the size of the system. The need for robust monitoring tools has become more evident with the advent of cloud…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-06-03 Jonathan Stuart Ward , Adam Barker