English
Related papers

Related papers: Scaling MPI Applications on Aurora

200 papers

We exploit floating-point DSPs in the Arria10 FPGA and multi-pumping feature of the M20K RAMs to build a dataflow-driven soft processor fabric for large graph workloads. In this paper, we introduce the idea of out-of-order node scheduling…

Hardware Architecture · Computer Science 2017-05-09 Siddhartha , Nachiket Kapre

We present a sparse linear system solver that is based on a multifrontal variant of Gaussian elimination, and exploits low-rank approximation of the resulting dense frontal matrices. We use hierarchically semiseparable (HSS) matrices, which…

Mathematical Software · Computer Science 2015-02-27 Pieter Ghysels , Xiaoye S. Li , Francois-Henry Rouet , Samuel Williams , Artem Napov

A considerable amount of research and engineering went into designing proxy applications, which represent common high-performance computing workloads, to co-design and evaluate the current generation of supercomputers, e.g., RIKEN's…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-04-18 Satoshi Matsuoka , Jens Domke , Mohamed Wahib , Aleksandr Drozd , Ray Bair , Andrew A. Chien , Jeffrey S. Vetter , John Shalf

Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-27 Luigi Fusco , Mikhail Khalilov , Marcin Chrapek , Giridhar Chukkapalli , Thomas Schulthess , Torsten Hoefler

Interconnection is crucial for computing systems. However, the current interconnection performance between processors and devices, such as memory devices and accelerators, significantly lags behind their computing performance, severely…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-02-21 Chen Chen , Xinkui Zhao , Guanjie Cheng , Yuesheng Xu , Shuiguang Deng , Jianwei Yin

Arm technology is becoming increasingly important in HPC. Recently, Fugaku, an \arm-based system, was awarded the number one place in the Top500 list. Raspberry Pis provide an inexpensive platform to become familiar with this architecture.…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-04-13 Nikunj Gupta , Steve R. Brandt , Bibek Wagle , Nanmiao , Alireza Kheirkhahan , Patrick Diehl , Hartmut Kaiser , Felix W. Baumann

The two main thrusts of computational science are more accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g.…

Computational Physics · Physics 2020-02-24 Saaketh Desai , Samuel Temple Reeve , James F. Belak

Autonomous robots require efficient on-device learning to adapt to new environments without cloud dependency. For this edge training, Microscaling (MX) data types offer a promising solution by combining integer and floating-point…

Hardware Architecture · Computer Science 2025-12-16 Stef Cuyckens , Xiaoling Yi , Nitish Satya Murthy , Chao Fang , Marian Verhelst

While many of the architectural details of future exascale-class high performance computer systems are still a matter of intense research, there appears to be a general consensus that they will be strongly heterogeneous, featuring…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-05 Moritz Kreutzer , Jonas Thies , Melven Röhrig-Zöllner , Andreas Pieper , Faisal Shahzad , Martin Galgon , Achim Basermann , Holger Fehske , Georg Hager , Gerhard Wellein

Reliable forecasts of the Earth system are crucial for human progress and safety from natural disasters. Artificial intelligence offers substantial potential to improve prediction accuracy and computational efficiency in this field, however…

Large-scale deep learning benefits from an emerging class of AI accelerators. Some of these accelerators' designs are general enough for compute-intensive applications beyond AI and Cloud TPU is one such example. In this paper, we…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-11-19 Kun Yang , Yi-Fan Chen , Georgios Roumpos , Chris Colby , John Anderson

It is currently not clear what the potential is of neuromorphic hardware beyond machine learning and neuroscience. In this project, a problem is investigated that is inherently difficult to fully implement in neuromorphic hardware by…

Neural and Evolutionary Computing · Computer Science 2019-12-02 Abdullahi Ali , Johan Kwisthout

Routing, switching, and the interconnect fabric are essential components in implementing large-scale neuromorphic computing architectures. While this fabric plays only a supporting role in the process of computing, for large AI workloads,…

Neural and Evolutionary Computing · Computer Science 2026-02-04 Madhuvanthi Srivatsav , Chiranjib Bhattacharyya , Shantanu Chakrabartty , Chetan Singh Thakur

Matrix multiplication is a foundational operation in scientific computing and machine learning, yet its computational complexity makes it a significant bottleneck for large-scale applications. The shift to parallel architectures, primarily…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-07-30 Mufakir Qamar Ansari , Mudabir Qamar Ansari

We explore the trade-offs of performing linear algebra using Apache Spark, compared to traditional C and MPI implementations on HPC platforms. Spark is designed for data analytics on cluster computing platforms with access to local disks…

This work brings the leading accuracy, sample efficiency, and robustness of deep equivariant neural networks to the extreme computational scale. This is achieved through a combination of innovative model architecture, massive…

Computational Physics · Physics 2023-04-21 Albert Musaelian , Anders Johansson , Simon Batzner , Boris Kozinsky

The unabated growth in AI workload demands is driving the need for concerted advances in compute, memory, and interconnect performance. As traditional semiconductor scaling slows, high-speed interconnects have emerged as the new scaling…

Hardware Architecture · Computer Science 2025-10-21 Mikhail Bernadskiy , Peter Carson , Thomas Graham , Taylor Groves , Ho John Lee , Eric Yeh

FPGA-based hardware accelerators have received increasing attention mainly due to their ability to accelerate deep pipelined applications, thus resulting in higher computational performance and energy efficiency. Nevertheless, the amount of…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-23 R. Nepomuceno , R. Sterle , G. Valarini , M. Pereira , H. Yviquel , G. Araujo

Interactive massively parallel computations are critical for machine learning and data analysis. These computations are a staple of the MIT Lincoln Laboratory Supercomputing Center (LLSC) and has required the LLSC to develop unique…

The resource demands of deep neural network (DNN) models introduce significant performance challenges, especially when deployed on resource-constrained edge devices. Existing solutions like model compression often sacrifice accuracy, while…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-26 Ziyang Zhang , Jie Liu , Luca Mottola