English
Related papers

Related papers: Scaling MPI Applications on Aurora

200 papers

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technologies include the Intel(TM) Xeon(TM) Data Center GPU Max…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-09 William E. Allcock , Benjamin S. Allen , James Anchell , Victor Anisimov , Thomas Applencourt , Abhishek Bagusetty , Ramesh Balakrishnan , Riccardo Balin , Solomon Bekele , Colleen Bertoni , Cyrus Blackworth , Renzo Bustamante , Kevin Canada , John Carrier , Christopher Chan-nui , Lance C. Cheney , Taylor Childers , Paul Coffman , Susan Coghlan , Tanima Dey , Michael D'Mello , Ashok Emani , Murali Emani , Kyle G. Felker , Sam Foreman , Olivier Franza , Longfei Gao , Marta García , María Garzarán , Balazs Gerofi , Yasaman Ghadar , Subrata Goswami , Neha Gupta , Kevin Harms , Väinö Hatanpää , Brian Holland , Carissa Holohan , Brian Homerding , Khalid Hossain , Xue Hu , Louise Huot , Huda Ibeid , Joseph A. Insley , Sai Jayanthi , Hong Jiang , Wei Jiang , Xiao-Yong Jin , Jeongnim Kim , Christopher Knight , Panagiotis Kourdis , Kalyan Kumaran , JaeHyuk Kwack , Janghaeng Lee , Ti Leggett , Ben Lenard , Chris Lewis , Nevin Liber , Johann Lombardi , Raymond M. Loy , Ye Luo , Bethany Lusch , Nilakantan Mahadevan , Beth Markey , Victor A. Mateevitsi , Gordon McPheeters , Ryan Milner , Jerome Mitchell , Vitali A. Morozov , Servesh Muralidharan , Tom Musta , Mrigendra Nagar , Vikram Narayana , Marieme Ngom , Anthony-Trung Nguyen , Nathan Nichols , Aditya Nishtala , James C. Osborn , Michael E. Papka , Scott Parker , Saumil S. Patel , Julia Piotrowska , Adrian C. Pope , Sucheta Raghunanda , Esteban Rangel , Paul M. Rich , Katherine M. Riley , Silvio Rizzi , Kris Rowe , Varuni Sastry , Adam Scovel , Filippo Simini , Haritha Siddabathuni Som , Patrick Steinbrecher , Rick Stevens , Xinmin Tian , Peter Upton , Thomas Uram , Archit K. Vasan , Álvaro Vázquez-Mayagoitia , Kaushik Velusamy , Brice Videau , Venkatram Vishwanath , Brian Whitney , Timothy J. Williams , Michael Woodacre , Sam Zeltner , Chuanjun Zhang , Gengbin Zheng , Huihuo Zheng

Sustaining exascale performance in production requires engineering choices and operational practices that emerge only under real deployment constraints and demand coordination across system layers. This paper reports experience from three…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-13 Kazushige Goto , Huda Ibeid , Kalyan Kumaran , Servesh Muralidharan , Anthony-Trung Nguyen , Aditya Nishtala

The Aurora supercomputer is an exascale-class system designed to tackle some of the most demanding computational workloads. Equipped with both High Bandwidth Memory (HBM) and DDR memory, it provides unique trade-offs in performance,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-07 Huda Ibeid , Vikram Narayana , Jeongnim Kim , Anthony Nguyen , Vitali Morozov , Ye Luo

As exascale systems reach unprecedented concurrency, traditional performance analysis tools struggle with the overhead of massive-scale telemetry. We present an accelerated infrastructure for the hpcanalysis framework that leverages a…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-12 Dragana Grbic

The objective of our research is to demonstrate the practical usage and orders of magnitude speedup of real-world applications by using alternative technologies to support high performance computing. Currently, the main barrier to the…

Astrophysics · Physics 2007-11-22 Robert J. Brunner , Volodymyr V. Kindratenko , Adam D. Myers

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-09-29 William F. Godoy , Pedro Valero-Lara , Caira Anderson , Katrina W. Lee , Ana Gainaru , Rafael Ferreira da Silva , Jeffrey S. Vetter

Niagara is currently the fastest supercomputer accessible to academics in Canada. It was deployed at the beginning of 2018 and has been serving the research community ever since. This homogeneous 60,000-core cluster, owned by the University…

With the announcement that the Aurora Supercomputer will be composed of general purpose Intel CPUs complemented by discrete high performance Intel GPUs, and the deployment of the oneAPI ecosystem, Intel has committed to enter the arena of…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-19 Yuhsiang M. Tsai , Terry Cojean , Hartwig Anzt

Production-quality parallel applications are often a mixture of diverse operations, such as computation- and communication-intensive, regular and irregular, tightly coupled and loosely linked operations. In conventional construction of…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-08-07 Ivy Bo Peng , Roberto Gioiosa , Gokcen Kestor , Erwin Laure , Stefano Markidis

The upcoming exascale computing systems Frontier and Aurora will draw much of their computing power from GPU accelerators. The hardware for these systems will be provided by AMD and Intel, respectively, each supporting their own GPU…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-05-18 Felix Wittwer , Nicholas K. Sauter , Derek Mendez , Billy K. Poon , Aaron S. Brewster , James M. Holton , Michael E. Wall , William E. Hart , Deborah J. Bard , Johannes P. Blaschke

AMD Instinct$^\text{TM}$ MI300A is the world's first data center accelerated processing unit (APU) with memory shared between the AMD "Zen 4" EPYC$^\text{TM}$ cores and third generation CDNA$^\text{TM}$ compute units. A single memory space…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-05-02 Suyash Tandon , Leopold Grinberg , Gheorghe-Teodor Bercea , Carlo Bertolli , Mark Olesen , Simone Bnà , Nicholas Malaya

With the advent of the Exascale capability allowing supercomputers to perform at least $10^{18}$ IEEE 754 Double Precision (64 bits) operations per second, many concerns have been raised regarding the energy consumption of high-performance…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-05-12 Tobias Fischbach , Emmanuel Kieffer , Pascal Bouvry

As dataset sizes increase, data analysis tasks in high performance computing (HPC) are increasingly dependent on sophisticated dataflows and out-of-core methods for efficient system utilization. In addition, as HPC systems grow, memory…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-01 George K. Thiruvathukal , Cameron Christensen , Xiaoyong Jin , François Tessier , Venkatram Vishwanath

Machine learning applications are computationally demanding and power intensive. Hardware acceleration of these software tools is a natural step being explored using various technologies. A recurrent processing unit (RPU) is fast and…

Emerging Technologies · Computer Science 2019-12-17 Heidi Komkov , Alessandro Restelli , Brian Hunt , Liam Shaughnessy , Itamar Shani , Daniel P. Lathrop

With at least 50 cores, Intel Xeon Phi is a true many-core architecture. Featuring fairly powerful cores, two cache levels, and very fast interconnections, the Xeon Phi can get a theoretical peak of 1000 GFLOPs and over 240 GB/s. These…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-12-23 Jianbin Fang , Ana Lucia Varbanescu , Henk Sips , Lilun Zhang , Yonggang Che , Chuanfu Xu

The Intel Xeon Phi manycore processor is designed to provide high performance matrix computations of the type often performed in data analysis. Common data analysis environments include Matlab, GNU Octave, Julia, Python, and R. Achieving…

We present direct astrophysical N-body simulations with up to a few million bodies using our parallel MPI/CUDA code on large GPU clusters in China, Ukraine and Germany, with different kinds of GPU hardware. These clusters are directly…

Instrumentation and Methods for Astrophysics · Physics 2013-12-09 P. Berczik , R. Spurzem , L. Wang , S. Zhong , O. Veles , I. Zinchenko , S. Huang , M. Tsai , G. Kennedy , S. Li , L. Naso , C. Li

In this paper, we analyze the performance and energy consumption of an Arm-based high-performance computing (HPC) system developed within the European project Mont-Blanc 3. This system, called Dibona, has been integrated by ATOS/Bull, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-13 Filippo Mantovani , Marta Garcia-Gasulla , José Gracia , Esteban Stafford , Fabio Banchelli , Marc Josep-Fabrego , Joel Criado-Ledesma , Mathias Nachtmann

With the rapidly growing demand for computing power new accelerator based architectures have entered the world of high performance computing since around 5 years. In particular GPGPUs have recently become very popular, however programming…

Performance · Computer Science 2013-08-16 Volker Weinberg , Momme Allalen

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-24 Harikrishna Tummalapalli , Riccardo Balin , Christine M. Simpson , Andrew Park , Aymen Alsaadi , Andrew E. Shao , Wesley Brewer , Shantenu Jha
‹ Prev 1 2 3 10 Next ›