English
Related papers

Related papers: Failure Data Analysis of HPC Systems

200 papers

In the current landscape of big data, the reliability and performance of storage systems are essential to the success of various applications and services. as data volumes continue to grow exponentially, the complexity and scale of the…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-02-06 Joshua Ludolf , Yesmin Reyna-Hernandez , Matthew Trevino

Software misconfiguration has consistently been a major reason for software failures. Over the past two decades, much work has been done to detect and diagnose software misconfigurations. However, there is still a gap between real-world…

Software Engineering · Computer Science 2026-05-29 Yuhao Liu , Yingnan Zhou , Hanfeng Zhang , Zhiwei Chang , Sihan Xu , Yan Jia , Wei Wang , Juncheng Hu , Zheli Liu

Fault-tolerance techniques depend on replication to enhance availability, albeit at the cost of increased infrastructure costs. This results in a fundamental trade-off: Fault-tolerant services must satisfy given availability and performance…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-06-16 Rasha Faqeh , Andrè Martin , Valerio Schiavoni , Pramod Bhatotia , Pascal Felber , Christof Fetzer

Resource demands of HPC applications vary significantly. However, it is common for HPC systems to primarily assign resources on a per-node basis to prevent interference from co-located workloads. This gap between the coarse-grained resource…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-03-14 Jie Li , George Michelogiannakis , Brandon Cook , Dulanya Cooray , Yong Chen

Blackouts in power grids typically result from cascading failures. The key importance of the electric power grid to society encourages further research into sustaining power system reliability and developing new methods to manage the risks…

Physics and Society · Physics 2014-02-18 Yakup Koç , Trivik Verma , Nuno A. M. Araujo , Martijn Warnier

Principal Component Analysis (PCA) is one of the most used tools for extracting low-dimensional representations of data, in particular for time series. Performances are known to strongly depend on the quality (amount of noise) and the…

Applications · Statistics 2024-12-16 Mariia Legenkaia , Laurent Bourdieu , Rémi Monasson

The first OpenFOAM HPC Challenge (OHC-1) was organised by the OpenFOAM HPC Technical Committee (HPCTC) to collect a snapshot of OpenFOAM's computational performance on contemporary production hardware and to compare hardware-constrained…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-31 Sergey Lesnik , Gregor Olenik , Mark Wassermann

Drug shortages occur frequently and are often caused by supply chain disruptions. For improvements to occur, it is necessary to be able to estimate the vulnerability of pharmaceutical supply chains. In this work, we present the first model…

Optimization and Control · Mathematics 2022-04-26 Emily L. Tucker , Mark S. Daskin

Silent Errors within hardware devices occur when an internal defect manifests in a part of the circuit which does not have check logic to detect the incorrect circuit operation. The results of such a defect can range from flipping a single…

Hardware Architecture · Computer Science 2022-03-18 Harish Dattatraya Dixit , Laura Boyle , Gautham Vunnam , Sneha Pendharkar , Matt Beadon , Sriram Sankar

Power consumption is a critical consideration in high performance computing systems and it is becoming the limiting factor to build and operate Petascale and Exascale systems. When studying the power consumption of existing systems running…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-01-27 Radim Vavřík , Antoni Portero , Štěpán Kuchař , Martin Golasowski , Simone Libutti , Giuseppe Massari , William Fornaciari , Vít Vondrák

Electrical power grids are vulnerable to cascading failures that can lead to large blackouts. Detection and prevention of cascading failures in power grids is impor- tant. Currently, grid operators mainly monitor the state (loading level)…

Physics and Society · Physics 2015-07-20 Martijn Warnier , Stefan Dulman , Yakup Koç , Eric Pauwels

Most FPGA boards in the HPC domain are well-suited for parallel scaling because of the direct integration of versatile and high-throughput network ports. However, the utilization of their network capabilities is often challenging and…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-09 Marius Meyer , Tobias Kenter , Lucian Petrica , Kenneth O'Brien , Michaela Blott , Christian Plessl

Cascading blackouts typically occur when nearly simultaneous outages occur in k out of N components in a power system, triggering subsequent failures that propagate through the network and cause significant load shedding. While large…

Computational Engineering, Finance, and Science · Computer Science 2019-04-12 Laurence A. Clarfeld , Paul D. H. Hines , Eric M. Hernandez , Margaret J. Eppstein

Too many defective compute chips are escaping existing manufacturing tests -- at least an order of magnitude more than industrial targets across all compute chip types in data centers. Silent data corruptions (SDCs) caused by test escapes,…

Network redundancy is one of the spatial network structural properties critical to robustness against cascading failures in power networks. The waiting-time distributions for network partitions in cascading failures explain how the spatial…

Physics and Society · Physics 2022-05-04 Long Huo , Xin Chen

The growing reliance on computer systems, particularly personal computers (PCs), necessitates heightened reliability to uphold user satisfaction. This research paper presents an in-depth analysis of extensive system telemetry data,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-02 Priyanka Mudgal , Rita H. Wouhaybi

Principal component analysis (PCA) is a most frequently used statistical tool in almost all branches of data science. However, like many other statistical tools, there is sometimes the risk of misuse or even abuse. In this paper, we…

Methodology · Statistics 2021-08-12 Xinyu Zhang , Howell Tong

Caching at the edge of wireless networks is a keytechnology to reduce traffic in the backhaul link. However, aconcentrated amount of requests during peak-periods may causethe outage of the system, meaning that the network is not ableto…

Networking and Internet Architecture · Computer Science 2021-01-05 Estefanía Recayte , Andrea Munari

Existing or planned power grids need to evaluate survivability under extreme events, like a number of peak load overloading conditions, which could possibly cause system collapses (i.e. blackouts). For realistic extreme events that are…

Systems and Control · Electrical Eng. & Systems 2026-03-13 Qinghua Ma , Reetam Sen Biswas , Denis Osipov , Guannan Qu , Soummya Kar , Shimiao Li

This study investigates the capabilities of Cyclic Redundancy Checks(CRCs) to detect burst and random errors. Researchers have favored these error detection codes throughout the evolution of computing and have implemented them in…

Networking and Internet Architecture · Computer Science 2022-05-24 Waylon Jepsen