English
Related papers

Related papers: MPCDF HPC Performance Monitoring System: Enabling …

200 papers

One of the more complex tasks for researchers using HPC systems is performance monitoring and tuning of their applications. Developing a practice of continuous performance improvement, both for speed-up and efficient use of resources is…

The ability to understand how a scientific application is executed on a large HPC system is of great importance in allocating resources within the HPC data center. In this paper, we describe how we used system performance data to identify:…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-09-15 David Brayford , Christoph Bernau , Wolfram Hesse , Carla Guillen

In this work, system monitoring and analysis are discussed in terms of their significance and benefits for operations and research in the field of high-performance computing (HPC). HPC systems deliver unique insights to computational…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-07-10 Florina M. Ciorba

System monitoring is an established tool to measure the utilization and health of HPC systems. Usually system monitoring infrastructures make no connection to job information and do not utilize hardware performance monitoring (HPM) data. To…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-02-19 Thomas Röhl , Jan Eitzinger , Georg Hager , Gerhard Wellein

High-performance computing (HPC) systems are a complex combination of software, processors, memory, networks, and storage systems characterized by frequent disruptive technological advances. Anomalous behavior has to be manually diagnosed…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-02-19 Charng-Da Lu

Workload characterization is an integral part of performance analysis of high performance computing (HPC) systems. An understanding of workload properties sheds light on resource utilization and can be used to inform performance…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-01-16 Nikolay A. Simakov , Joseph P. White , Robert L. DeLeon , Steven M. Gallo , Matthew D. Jones , Jeffrey T. Palmer , Benjamin Plessinger , Thomas R. Furlani

Performance analysis is an essential task in High-Performance Computing (HPC) systems and it is applied for different purposes such as anomaly detection, optimal resource allocation, and budget planning. HPC monitoring tasks generate a huge…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-12-12 Mohamed S. Halawa , Rebeca P. Díaz-Redondo , Ana Fernández-Vilas

High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPIs), such as resource…

Artificial Intelligence · Computer Science 2023-12-12 Mohamed Soliman Halawa , Rebeca P. Díaz-Redondo , Ana Fernández-Vilas

The increasing use and cost of high performance computing (HPC) requires new easy-to-use tools to enable HPC users and HPC systems engineers to transparently understand the utilization of resources. The MIT Lincoln Laboratory Supercomputing…

In high-performance computing (HPC) environments, system monitoring data is often unlabeled and high-dimensional, making it difficult to reliably detect and understand anomalous computing nodes. The growing scale and dimensionality of the…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-15 Allison Austin , Shilpika , Yan To Linus Lam , Yun-Hsin Kuo , Venkatram Vishwanath , Michael E. Papka , Kwan-Liu Ma

Many tools and libraries employ hardware performance monitoring (HPM) on modern processors, and using this data for performance assessment and as a starting point for code optimizations is very popular. However, such data is only useful if…

Performance · Computer Science 2013-02-20 Jan Treibig , Georg Hager , Gerhard Wellein

The ability to monitor and interpret of hardware system events and behaviors are crucial to improving the robustness and reliability of these systems, especially in a supercomputing facility. The growing complexity and scale of these…

Human-Computer Interaction · Computer Science 2023-06-19 Shilpika , Bethany Lusch , Murali Emani , Filippo Simini , Venkatram Vishwanath , Michael E. Papka , Kwan-Liu Ma

In recent years, several HPC facilities have started continuous monitoring of their systems and jobs to collect performance-related data for understanding performance and operational efficiency. Such data can be used to optimize the…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-03 Ian J. Costello , Abhinav Bhatele

We present a slow control system to gather all relevant environment information necessary to effectively and reliably run an HPC (High Performance Computing) system at a high value over price ratio. The scalable and reliable overall concept…

Other Computer Science · Computer Science 2018-02-05 Peter Bernd Otte , Dalibor Djukanovic

Today's HPC installations are highly-complex systems, and their complexity will only increase as we move to exascale and beyond. At each layer, from facilities to systems, from runtimes to applications, a wide range of tuning decisions must…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-08-15 Alessio Netti , Micha Mueller , Axel Auweter , Carla Guillen , Michael Ott , Daniele Tafani , Martin Schulz

Energy-centric design is paramount in the current embedded computing era: use cases require increasingly high performance at an affordable power budget, often under real-time constraints. Hardware heterogeneity and parallelism help address…

Performance · Computer Science 2025-07-01 Sergio Mazzola , Gabriele Ara , Thomas Benz , Björn Forsberg , Tommaso Cucinotta , Luca Benini

Computer-based scientific experiments are becoming increasingly data-intensive, necessitating the use of High-Performance Computing (HPC) clusters to handle large scientific workflows. These workflows result in complex data and control…

Databases · Computer Science 2025-02-17 Zahra Sadeghibogar , Alessandro Berti , Marco Pegoraro , Wil M. P. van der Aalst

This paper addresses the design of an event-triggered, data-based, and performance-oriented adaption method for model predictive control (MPC). The performance of such a strategy strongly depends on the accuracy of the prediction model,…

Systems and Control · Electrical Eng. & Systems 2026-03-13 Samuel Mallick , Laura Boca de de Giuli , Alessio La Bella , Azita Dabiri , Bart De Schutter , Riccardo Scattolini

Running scientific workflows on a supercomputer can be a daunting task for a scientific domain specialist. Workflow management solutions (WMS) are a standard method for reducing the complexity of application deployment on high performance…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-07-30 Wouter Klijn , Sandra Diaz-Pier , Abigail Morrison , Alexander Peyser
‹ Prev 1 2 3 10 Next ›