English
Related papers

Related papers: Failure Data Analysis of HPC Systems

200 papers

Today's high-performance computing (HPC) systems are heavily instrumented, generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-08-24 Byung H. Park , Saurabh Hukerikar , Ryan Adamson , Christian Engelmann

Power grid outages cause huge economical and societal costs. Disruptions in the power distribution grid are responsible for a significant fraction of electric power unavailability to customers. The impact of extreme weather conditions,…

Systems and Control · Computer Science 2015-06-30 Yakup Koc , Abhishek Raman , Martijn Warnier , Tarun Kumar

Software reliability models are an important tool in quality management and release planning. There is a large number of different models that often exhibit strengths in different areas. This paper proposes a model that is based on a…

Software Engineering · Computer Science 2016-12-13 Stefan Wagner , Helmut Fischer

Modern power networks face increasing vulnerability to cascading failures due to high complexity and the growing penetration of intermittent resources, necessitating rigorous security assessment beyond the conventional $N-1$ criterion.…

Systems and Control · Electrical Eng. & Systems 2025-12-10 Alexandre Gracia-Calvo , Francesca Rossi , Eduardo Iraola , Juan Carlos Olives-Camps , Eduardo Prieto-Araujo

We tackle the challenging tasks of monitoring on unstable HPC platforms the performance of CFD applications all along their development. We have designed and implemented a monitoring framework, integrated at the end of a CI-CD pipeline.…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-01-17 Damien Dosimont , Guillaume Houzeaux

The rapid development of large clusters built with commodity hardware has highlighted scalability issues with deploying and effectively running system software in large clusters. We describe here our experiences with monitoring, image…

Computational Physics · Physics 2007-05-23 A. Chan , R. Hogue , C. Hollowell , O. Rind , T. Throwe , T. Wlodek

Silent data corruption (SDC) threatens the reliability of large-scale GPU clusters used for training large language models, yet its rarity and lack of explicit error signals make accurate high-level modeling challenging. To address this…

Modern software systems become too complex to be tested and validated. Detecting software partial failures in complex systems at runtime assist to handle software unintended behaviors, avoiding catastrophic software failures and improving…

Software Engineering · Computer Science 2021-10-14 Shiyi Kong , Minyan Lu , Jun Ai , Shuguang Wang

Automated driving functions at high levels of autonomy operate without driver supervision. The system itself must provide suitable responses in case of hardware element failures. This requires fault-tolerant approaches using domain ECUs and…

Computers and Society · Computer Science 2022-10-11 Tim Maurice Julitz , Antoine Tordeux , Manuel Löwer

The dataset was collected for 332 compute nodes throughout May 19 - 23, 2023. May 19 - 22 characterizes normal compute cluster behavior, while May 23 includes an anomalous event. The dataset includes eight CPU, 11 disk, 47 memory, and 22…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-11-29 Diana McSpadden , Yasir Alanazi , Bryan Hess , Laura Hild , Mark Jones , Yiyang Lub , Ahmed Mohammed , Wesley Moore , Jie Ren , Malachi Schram , Evgenia Smirni

This paper studies the stability and $\mathcal{H}_{\infty}$ performance analysis problem for linear networked and quantized control systems with both communication delays random packet losses. To deal with the network-induced uncertainties…

Systems and Control · Electrical Eng. & Systems 2021-03-05 Wei Ren , Junlin Xiong

Modern grid monitoring equipment enables utilities to collect detailed records of power interruptions. These data are aggregated to compute publicly reported metrics describing high-level characteristics of grid performance. The current…

Physics and Society · Physics 2019-01-14 Laurel N. Dunn , Michael D. Sohn , Kristina Hamachi Lacommare , Joseph H. Eto

Automated Program Repair (APR) techniques have drawn wide attention from both academia and industry. Meanwhile, one main limitation with the current state-of-the-art APR tools is that patches passing all the original tests are not…

Software Engineering · Computer Science 2023-02-10 Jun Yang , Yuehan Wang , Yiling Lou , Ming Wen , Lingming Zhang

Power outages have become increasingly frequent, intense, and prolonged in the US due to climate change, aging electrical grids, and rising energy demand. However, largely due to the absence of granular spatiotemporal outage data, we lack…

Computers and Society · Computer Science 2024-10-30 Junwei Ma , Bo Li , Olufemi A. Omitaomu , Ali Mostafavi

This paper considers the stability analysis for nonlinear sampled-data systems with failures in the feedback loop. The failures are caused by shared resources, and modeled by a weakly hard real-time (WHRT) dropout description. The WHRT…

Systems and Control · Electrical Eng. & Systems 2020-05-18 Michael Hertneck , Steffen Linsenmayer , Frank Allgöwer

Migrating heterogeneous high-performance computing (HPC) systems to resource-aware scheduling introduces both technical and behavioral challenges, particularly in production environments with established user workflows. This paper presents…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-31 Glen MacLachlan , Joseph Creech , Rubeel Muhammad Iqbal , Clark Gaylord , Jake Messick

Health monitoring, fault analysis, and detection are critical for the safe and sustainable operation of battery systems. We apply Gaussian process resistance models on lithium iron phosphate battery field data to effectively separate the…

Machine Learning · Computer Science 2024-10-31 Joachim Schaeffer , Eric Lenz , Duncan Gulla , Martin Z. Bazant , Richard D. Braatz , Rolf Findeisen

In recent years, several HPC facilities have started continuous monitoring of their systems and jobs to collect performance-related data for understanding performance and operational efficiency. Such data can be used to optimize the…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-03 Ian J. Costello , Abhinav Bhatele

Maintenance is the last and the most critical phase of the software development life cycle. It involves debugging of errors and different types of enhancements which are requested by the user. Software reliability regarding maintenance is…

Software Engineering · Computer Science 2016-05-04 Ahmad Mateen , Muhammad Azeem Akbar

The state-of-the-art approach to manage blockchains is to process blocks of transactions in a shared-nothing environment. Although blockchains have the potential to provide various services for high-performance computing (HPC) systems, HPC…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-22 Abdullah Al-Mamun , Dongfang Zhao
‹ Prev 1 8 9 10 Next ›