English
Related papers

Related papers: Towards Adaptive Resilience in High Performance Co…

200 papers

Orchestrated collaborative effort of physical and cyber components to satisfy given requirements is the central concept behind Cyber-Physical Systems (CPS). To duly ensure the performance of components, a software-based resilience manager…

Software Engineering · Computer Science 2020-04-14 Mohammad Shihabul Haque , Daniel Jun Xian Ng , Arvind Easwaran , Karthik Thangamariappan

Dynamic resource management opens up numerous opportunities in High Performance Computing. It improves the system-level services as well as application performance. Checkpointing can also be deemed as a system-level service and can reap the…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-11-09 Jophin John , Michael Gerndt

As high-performance computing systems scale in size and computational power, the danger of silent errors, i.e., errors that can bypass hardware detection mechanisms and impact application state, grows dramatically. Consequently,…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-09-06 Luanzheng Guo , Dong Li , Ignacio Laguna , Martin Schulz

Over the past decade, extreme weather events have significantly increased worldwide, leading to widespread power outages and blackouts. As these threats continue to challenge power distribution systems, the importance of mitigating the…

Systems and Control · Electrical Eng. & Systems 2023-08-16 Shuva Paul , Abodh Poudyal , Shiva Poudel , Anamika Dubey , Zhaoyu Wang

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High- Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model,…

Deep learning (DL) applications are increasingly being deployed on HPC systems, to leverage the massive parallelism and computing power of those systems for DL model training. While significant effort has been put to facilitate distributed…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-30 Elvis Rojas , Albert Njoroge Kahira , Esteban Meneses , Leonardo Bautista Gomez , Rosa M Badia

The growing energy demands of HPC systems have made energy efficiency a critical concern for system developers and operators. However, HPC users are generally less aware of how these energy concerns influence the design, deployment, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-28 Estela Suarez , Jorge Amaya , Martin Frank , Oliver Freyermuth , Maria Girone , Bartosz Kostrzewa , Susanne Pfalzner

High Performance Computing (HPC) aims at providing reasonably fast computing solutions to scientific and real life problems. The advent of multicore architectures is noticeable in the HPC history, because it has brought the underlying…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-07 Claude Tadonki

Volunteer Computing, sometimes called Public Resource Computing, is an emerging computational model that is very suitable for work-pooled parallel processing. As more complex grid applications make use of work flows in their design and…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-11-27 Lei Ni , Aaron Harwood

Reliability is a cumbersome problem in High Performance Computing Systems and Data Centers evolution. During operation, several types of fault conditions or anomalies can arise, ranging from malfunctioning hardware to improper…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-30 Andrea Borghesi , Antonio Libri , Luca Benini , Andrea Bartolini

Health monitoring applications increasingly rely on machine learning techniques to learn end-user physiological and behavioral patterns in everyday settings. Considering the significant role of wearable devices in monitoring human body…

Machine Learning · Computer Science 2022-08-03 Sina Shahhosseini , Yang Ni , Hamidreza Alikhani , Emad Kasaeyan Naeini , Mohsen Imani , Nikil Dutt , Amir M. Rahmani

Recovery from transient failures is one of the prime issues in the context of distributed systems. These systems demand to have transparent yet efficient techniques to achieve the same. Checkpoint is defined as a designated place in a…

Networking and Internet Architecture · Computer Science 2011-09-01 Ruchi Tuli , Parveen Kumar

As High-Performance Computing (HPC) systems strive towards the exascale goal, studies suggest that they will experience excessive failure rates. For this reason, detecting and classifying faults in HPC systems as they occur and initiating…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-07-12 Alessio Netti , Zeynep Kiziltan , Ozalp Babaoglu , Alina Sirbu , Andrea Bartolini , Andrea Borghesi

Industrial Internet of Things (I-IoT) enables fully automated production systems by continuously monitoring devices and analyzing collected data. Machine learning methods are commonly utilized for data analytics in such systems.…

Cryptography and Security · Computer Science 2022-03-17 Onat Gungor , Tajana Rosing , Baris Aksanli

This paper presents a robust adaptive learning Model Predictive Control (MPC) framework for linear systems with parametric uncertainties and additive disturbances performing iterative tasks. The approach refines the parameter estimates…

Systems and Control · Electrical Eng. & Systems 2025-09-04 Hannes Petrenz , Johannes Köhler , Francesco Borrelli

Collective adaptive systems are an emerging class of networked computational systems, particularly suited in application domains such as smart cities, complex sensor networks, and the Internet of Things. These systems tend to feature large…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-11-23 Mirko Viroli , Giorgio Audrito , Jacob Beal , Ferruccio Damiani , Danilo Pianini

We propose a novel framework for designing a resilient Model Predictive Control (MPC) targeting uncertain linear systems under cyber attack. Assuming a periodic attack scenario, we model the system under Denial of Service (DoS) attack, also…

Systems and Control · Electrical Eng. & Systems 2023-10-16 Milad Farsi , Shuhao Bian , Nasser L. Azad , Xiaobing Shi , Andrew Walenstein

High intensive computation applications can usually take days to months to finish an execution. During this time, it is common to have variations of the available resources when considering that such hardware is usually shared among a…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-01-27 Kiran Mantripragada , Alecio Binotto , Leonardo P. Tizzei

In most of modern enterprise systems, redundancy configuration is often considered to provide availability during the part of such systems is being patched. However, the redundancy may increase the attack surface of the system. In this…

Cryptography and Security · Computer Science 2017-05-02 Mengmeng Ge , Huy Kang Kim , Dong Seong Kim

High Performance Computing (HPC) is a highly demanded discipline in companies and institutions. However, as students and also afterwards as professors, we observed a lack of HPC related content in the engineering degrees at our university,…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-05 S. Catalán , R. Carratalá-Sáez , S. Iserte