English
Related papers

Related papers: Performance Analysis of HPC applications on the Au…

200 papers

Next-generation supercomputers will feature more hierarchical and heterogeneous memory systems with different memory technologies working side-by-side. A critical question is whether at large scale existing HPC applications and emerging…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-04-27 Ivy Bo Peng , Stefano Markidis , Erwin Laure , Gokcen Kestor , Roberto Gioiosa

The Aurora supercomputer, which was deployed at Argonne National Laboratory in 2024, is currently one of three Exascale machines in the world on the Top500 list. The Aurora system is composed of over ten thousand nodes each of which…

Many high end and next generation computing systems to incorporated alternative memory technologies to meet performance goals. Since these technologies present distinct advantages and tradeoffs compared to conventional DDR* SDRAM, such as…

Performance · Computer Science 2021-10-06 M. Ben Olson , Brandon Kammerdiener , Kshitij A. Doshi , Terry Jones , Michael R. Jantz

Discrete GPUs are a cornerstone of HPC and data center systems, requiring management of separate CPU and GPU memory spaces. Unified Virtual Memory (UVM) has been proposed to ease the burden of memory management; however, at a high cost in…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-14 Jacob Wahlgren , Gabin Schieffer , Ruimin Shi , Edgar A. León , Roger Pearce , Maya Gokhale , Ivy Peng

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technologies include the Intel(TM) Xeon(TM) Data Center GPU Max…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-09 William E. Allcock , Benjamin S. Allen , James Anchell , Victor Anisimov , Thomas Applencourt , Abhishek Bagusetty , Ramesh Balakrishnan , Riccardo Balin , Solomon Bekele , Colleen Bertoni , Cyrus Blackworth , Renzo Bustamante , Kevin Canada , John Carrier , Christopher Chan-nui , Lance C. Cheney , Taylor Childers , Paul Coffman , Susan Coghlan , Tanima Dey , Michael D'Mello , Ashok Emani , Murali Emani , Kyle G. Felker , Sam Foreman , Olivier Franza , Longfei Gao , Marta García , María Garzarán , Balazs Gerofi , Yasaman Ghadar , Subrata Goswami , Neha Gupta , Kevin Harms , Väinö Hatanpää , Brian Holland , Carissa Holohan , Brian Homerding , Khalid Hossain , Xue Hu , Louise Huot , Huda Ibeid , Joseph A. Insley , Sai Jayanthi , Hong Jiang , Wei Jiang , Xiao-Yong Jin , Jeongnim Kim , Christopher Knight , Panagiotis Kourdis , Kalyan Kumaran , JaeHyuk Kwack , Janghaeng Lee , Ti Leggett , Ben Lenard , Chris Lewis , Nevin Liber , Johann Lombardi , Raymond M. Loy , Ye Luo , Bethany Lusch , Nilakantan Mahadevan , Beth Markey , Victor A. Mateevitsi , Gordon McPheeters , Ryan Milner , Jerome Mitchell , Vitali A. Morozov , Servesh Muralidharan , Tom Musta , Mrigendra Nagar , Vikram Narayana , Marieme Ngom , Anthony-Trung Nguyen , Nathan Nichols , Aditya Nishtala , James C. Osborn , Michael E. Papka , Scott Parker , Saumil S. Patel , Julia Piotrowska , Adrian C. Pope , Sucheta Raghunanda , Esteban Rangel , Paul M. Rich , Katherine M. Riley , Silvio Rizzi , Kris Rowe , Varuni Sastry , Adam Scovel , Filippo Simini , Haritha Siddabathuni Som , Patrick Steinbrecher , Rick Stevens , Xinmin Tian , Peter Upton , Thomas Uram , Archit K. Vasan , Álvaro Vázquez-Mayagoitia , Kaushik Velusamy , Brice Videau , Venkatram Vishwanath , Brian Whitney , Timothy J. Williams , Michael Woodacre , Sam Zeltner , Chuanjun Zhang , Gengbin Zheng , Huihuo Zheng

In the rapidly evolving domain of high-performance computing (HPC), heterogeneous architectures such as the SX-Aurora TSUBASA (SX-AT) system architecture, which integrate diverse processor types, present both opportunities and challenges…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-25 Shubham , Keichi Takahashi , Hiroyuki Takizawa

In this paper we explore the performance of Intel Xeon MAX CPU Series, representing the most significant new variation upon the classical CPU architecture since the Intel Xeon Phi Processor. Given the availability of a large on-package…

Performance · Computer Science 2023-09-19 Istvan Z Reguly

Hardware accelerators have become a de-facto standard to achieve high performance on current supercomputers and there are indications that this trend will increase in the future. Modern accelerators feature high-bandwidth memory next to the…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-06-07 Ivy Bo Peng , Roberto Gioiosa , Gokcen Kestor , Erwin Laure , Stefano Markidis

Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an essential step toward achieving efficient and reliable HPC…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-05 Zhu Zhu , Yu Sun , Dhatri Parakal , Bo Fang , Steven Farrell , Gregory H. Bauer , Brett Bode , Ian T. Foster , Michael E. Papka , William Gropp , Zhao Zhang , Lishan Yang

Cutting-edge embedded system applications, such as self-driving cars and unmanned drone software, are reliant on integrated CPU/GPU platforms for their DNNs-driven workload, such as perception and other highly parallel components. In this…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-03-20 Soroush Bateni , Zhendong Wang , Yuankun Zhu , Yang Hu , Cong Liu

Byte-addressable persistent memory (B-APM) presents a new opportunity to bridge the performance gap between main memory and storage. In this paper, we present the usage scenarios for this new technology, based on the capabilities of Intel's…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-12-14 Michele Weiland , Bernhard Homoelle

As high-performance computing (HPC) moves into the exascale era, computer scientists and engineers must find innovative ways of transferring and processing unprecedented amounts of data. As the scale and complexity of the applications…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-09-30 Melissa Romanus , Robert B. Ross , Manish Parashar

Current HPC systems provide memory resources that are statically configured and tightly coupled with compute nodes. However, workloads on HPC systems are evolving. Diverse workloads lead to a need for configurable memory resources to…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-03-23 Jacob Wahlgren , Maya Gokhale , Ivy B. Peng

Somewhat surprisingly, the behavior of analytical query engines is crucially affected by the dynamic memory allocator used. Memory allocators highly influence performance, scalability, memory efficiency and memory fairness to other…

Databases · Computer Science 2019-05-20 Dominik Durner , Viktor Leis , Thomas Neumann

The latest trends in high-performance computing systems show an increasing demand on the use of a large scale multicore systems in a efficient way, so that high compute-intensive applications can be executed reasonably well. However, the…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-02-25 Juliana M. N. Silva , Cristina Boeres , Lúcia M. A. Drummond , Artur A. Pessoa

Sustaining exascale performance in production requires engineering choices and operational practices that emerge only under real deployment constraints and demand coordination across system layers. This paper reports experience from three…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-13 Kazushige Goto , Huda Ibeid , Kalyan Kumaran , Servesh Muralidharan , Anthony-Trung Nguyen , Aditya Nishtala

Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-21 Nan Ding , Pieter Maris , Hai Ah Nam , Taylor Groves , Muaaz Gul Awan , LeAnn Lindsey , Christopher Daley , Oguz Selvitopi , Leonid Oliker , Nicholas Wright , Samuel Williams

Multicore CPU architectures have been established as a structure for general-purpose systems for high-performance processing of applications. Recent multicore CPU has evolved as a system architecture based on non-uniform memory…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-01-26 Geunsik Lim , Sang-Bum Suh

We explored the possible benefits of integrating quantum simulators in a "hybrid" quantum machine learning (QML) workflow that uses both classical and quantum computations in a high-performance computing (HPC) environment. Here, we used two…

Emerging Technologies · Computer Science 2024-07-11 Samuel T. Bieberich , Michael A. Sandoval

High-Performance Computing (HPC) and Artificial Intelligence (AI) workloads typically demand substantial memory bandwidth and, to a degree, memory capacity. CXL memory expansion modules, also known as CXL "type-3" devices, enable…

Operating Systems · Computer Science 2024-12-18 Rohit Sehgal , Vishal Tanna , Vinicius Petrucci , Anil Godbole
‹ Prev 1 2 3 10 Next ›