English
Related papers

Related papers: Failure Data Analysis of HPC Systems

200 papers

Server Availability (SA) is an important measure of overall systems security. Important security systems rely on the availability of their hosting servers to deliver critical security services. Many of these servers offer management…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-01-23 Ayman M. Bahaa-Eldin , Hoda K. Mohamead , Sally S. Deraz

We discuss some specific software engineering challenges in the field of high-performance computing, and argue that the slow adoption of SE tools and techniques is at least in part caused by the fact that these do not address the HPC…

Software Engineering · Computer Science 2021-12-14 Jonas Thies , Melven Röhrig-Zöllner , Achim Basermann

We present ANSC, a probabilistic capacity health scoring framework for hyperscale datacenter fabrics. While existing alerting systems detect individual device or link failures, they do not capture the aggregate risk of cascading capacity…

Networking and Internet Architecture · Computer Science 2025-08-25 Madhava Gaikwad , Abhishek Gandhi

Software upgrades are critical to maintaining server reliability in datacenters. While job duration prediction and scheduling have been extensively studied, the unique challenges posed by software upgrades remain largely under-explored.…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-19 Yi Ding , Aijia Gao , Thibaud Ryden , Michal Sedlak , Essam Ewaisha , Igor Marnat , Henry Hoffmann

In today's corporate landscape, particularly where operations rely heavily on information technologies, establishing a robust business continuity plan, including a disaster recovery strategy, is essential for ensuring swift recuperation…

Networking and Internet Architecture · Computer Science 2025-12-02 Saso Nikolovski , Pece Mitrevski

As multimodal and AI-driven services exchange hundreds of megabytes per request, existing IPC runtimes spend a growing share of CPU cycles on memory copies. Although both hardware and software mechanisms are exploring memory offloading,…

Operating Systems · Computer Science 2026-01-13 Misun Park , Richi Dubey , Yifan Yuan , Nam Sung Kim , Ada Gavrilovska

The rise of AI and the economic dominance of cloud computing have created a new nexus of innovation for high performance computing (HPC), which has a long history of driving scientific discovery. In addition to performance needs, scientific…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-10 Vanessa Sochat , Daniel Milroy , Abhik Sarkar , Aniruddha Marathe , Tapasya Patki

Achieving connectivity reliability is one of the significant challenges for 5G and beyond 5G cellular networks. The present understanding of reliability in the context of mobile communication does not adequately cover the stochastic…

Networking and Internet Architecture · Computer Science 2024-09-04 Subhyal Bin Iqbal , Behnam Khodapanah , Philipp Schulz , Gerhard P. Fettweis

The total estimated energy bill for data centers in 2010 was \$11.5 billion, and experts estimate that the energy cost of a typical data center doubles every five years. On the other hand, computational developments have started to lag…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-12-01 Álvaro García-Recuero

Driven by artificial intelligence, data science, and high-resolution simulations, I/O workloads and hardware on high-performance computing (HPC) systems have become increasingly complex. This complexity can lead to large I/O overheads and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-03 Hammad Ather , Jean Luca Bez , Chen Wang , Hank Childs , Allen D. Malony , Suren Byna

All computing infrastructure suffers from performance variability, be it bare-metal or virtualized. This phenomenon originates from many sources: some transient, such as noisy neighbors, and others more permanent but sudden, such as changes…

Performance · Computer Science 2020-03-11 Dmitry Duplyakin , Alexandru Uta , Aleksander Maricq , Robert Ricci

The use of learning-based techniques to achieve automated software vulnerability detection has been of longstanding interest within the software security domain. These data-driven solutions are enabled by large software vulnerability…

Software Engineering · Computer Science 2023-01-16 Roland Croft , M. Ali Babar , Mehdi Kholoosi

Large-scale power failures are induced by nearly all natural disasters from hurricanes to wild fires. A fundamental problem is whether and how recovery guided by government policies is able to meet the challenge of a wide range of…

Computational Engineering, Finance, and Science · Computer Science 2021-01-01 Amir Hossein Afsharinejad , Chuanyi Ji , Robert Wilcox

Modern operating systems are typically POSIX-compliant with major system calls specified decades ago. The next generation of non-volatile memory (NVM) technologies raise concerns about the efficiency of the traditional POSIX-based systems.…

Operating Systems · Computer Science 2019-03-12 Viacheslav Dubeyko , Om Rameshwar Gatla , Mai Zheng

This paper studies the robustness of observability of a linear time-invariant system under sensor failures from a computational perspective. To be precise, the problem of determining the minimum number of sensors whose removal can destroy…

Optimization and Control · Mathematics 2023-07-18 Yuan Zhang , Yuanqing Xia , Kun Liu

The delivery of key services in domains ranging from finance and manufacturing to healthcare and transportation is underpinned by a rapidly growing number of mission-critical enterprise applications. Ensuring the continuity of these complex…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-23 Premathas Somasekaram , Radu Calinescu , Rajkumar Buyya

Time delays are a common perturbation in systems with many states, such as networked, distributed, or decentralized systems. Current methods analyzing the stability of large systems with time delay typically produce very conservative…

Systems and Control · Computer Science 2017-10-31 George Armanious , Rick Lind

Archiving and systematic backup of large digital data generates a quick demand for multi-peta byte scale storage systems. As drive capacities continue to grow beyond the few terabytes range to address the demands of today's cloud, the…

Information Theory · Computer Science 2018-10-26 Suayb S. Arslan

The mean time between failures (MTBF) of HPC systems is rapidly reducing, and that current failure recovery mechanisms e.g., checkpoint-restart, will no longer be able to recover the systems from failures. Early failure detection is a new…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-03-15 Siavash Ghiasvand , Florina M. Ciorba , Wolfgang E. Nagel

Cascading failures in power systems normally occur as a result of initial disturbance or faults on electrical elements, closely followed by errors of human operators. It remains a great challenge to systematically trace the source of…

Systems and Control · Computer Science 2017-03-16 Chao Zhai , Hehong Zhang , Gaoxi Xiao , Tso-Chien Pan