Power-Capping Metric Evaluation for Improving Energy Efficiency in HPC Applications
Distributed, Parallel, and Cluster Computing2025-06-26v2Computational Engineering, Finance, and SciencePerformanceSystems and ControlSystems and Control
With high-performance computing systems now running at exascale, optimizing power-scaling management and resource utilization has become more critical than ever. This paper explores runtime power-capping optimizations that leverage integrated CPU-GPU power management on architectures like the NVIDIA GH200 superchip. We evaluate energy-performance metrics that account for simultaneous CPU and GPU power-capping effects by using two complementary approaches: speedup-energy-delay and a Euclidean distance-based multi-objective optimization method. By targeting a mostly compute-bound exascale science application, the Locally Self-Consistent Multiple Scattering (LSMS), we explore challenging scenarios to identify potential opportunities for energy savings in exascale applications, and we recognize that even modest reductions in energy consumption can have significant overall impacts. Our results highlight how GPU task-specific dynamic power-cap adjustments combined with integrated CPU-GPU power steering can improve the energy utilization of certain GPU tasks, thereby laying the groundwork for future adaptive optimization strategies.
@article{arxiv.2505.21758,
title = {Power-Capping Metric Evaluation for Improving Energy Efficiency in HPC Applications},
author = {Maria Patrou and Thomas Wang and Wael Elwasif and Markus Eisenbach and Ross Miller and William Godoy and Oscar Hernandez},
journal= {arXiv preprint arXiv:2505.21758},
year = {2025}
}
Comments
14 pages, 3 figures, 2 tables. Accepted at the Energy Efficiency with Sustainable Performance: Techniques, Tools, and Best Practices, EESP Workshop, in conjunction with ISC High Performance 2025