English
Related papers

Related papers: Performance Evaluation of a Next-Generation SX-Aur…

200 papers

The SX-Aurora TSUBASA PCIe accelerator card is the newest model of NEC's SX architecture family. Its multi-core vector processor features a vector length of 16 kbits and interfaces with up to 48 GB of HBM2 memory in the current models,…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-02-04 Benjamin Huth , Nils Meyer , Tilo Wettig

In the rapidly evolving domain of high-performance computing (HPC), heterogeneous architectures such as the SX-Aurora TSUBASA (SX-AT) system architecture, which integrate diverse processor types, present both opportunities and challenges…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-25 Shubham , Keichi Takahashi , Hiroyuki Takizawa

We have proposed a method to accelerate the computation of Kubo formula optimized to vector processors. The key concept is parallel evaluation of multiple integration points, enabled by batched linear algebra operations. Through benchmark…

Materials Science · Physics 2023-09-14 Yuta Yahagi , Toshihiro Kato

Vector architectures are essential for boosting computing throughput. ARM provides SVE as the next-generation length-agnostic vector extension beyond traditional fixed-length SIMD. This work provides a first study of the maturity and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-15 Ruimin Shi , Gabin Schieffer , Maya Gokhale , Pei-Hung Lin , Hiren Patel , Ivy Peng

The Aurora supercomputer is an exascale-class system designed to tackle some of the most demanding computational workloads. Equipped with both High Bandwidth Memory (HBM) and DDR memory, it provides unique trade-offs in performance,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-07 Huda Ibeid , Vikram Narayana , Jeongnim Kim , Anthony Nguyen , Vitali Morozov , Ye Luo

This living paper reviews the present High Performance Computing (HPC) capabilities of the Tinker-HP molecular modeling package. We focus here on the reference, double precision, massively parallel molecular dynamics engine present in…

Mathematical Software · Computer Science 2024-01-11 Luc-Henri Jolly , Alejandro Duran , Louis Lagardère , Jay W. Ponder , Pengyu Ren , Jean-Philip Piquemal

Vector architectures are gaining traction for highly efficient processing of data-parallel workloads, driven by all major ISAs (RISC-V, Arm, Intel), and boosted by landmark chips, like the Arm SVE-based Fujitsu A64FX, powering the TOP500…

Hardware Architecture · Computer Science 2025-01-10 Matteo Perotti , Matheus Cavalcante , Nils Wistoff , Renzo Andri , Lukas Cavigelli , Luca Benini

High Performance Computing (HPC) platforms allow scientists to model computationally intensive algorithms. HPC clusters increasingly use General-Purpose Graphics Processing Units (GPGPUs) as accelerators; FPGAs provide an attractive…

Hardware Architecture · Computer Science 2015-04-20 Syed Waqar Nabi , Saji N. Hameed , Wim Vanderbauwhede

The performance of the Hybrid Monte Carlo algorithm is determined by the speed of sparse matrix-vector multiplication within the context of preconditioned conjugate gradient iteration. We study these operations as implemented for the…

Statistical Mechanics · Physics 2016-08-14 Kyle A. Wendt , Joaquín E. Drut , Timo A. Lähde

The emergence of novel hardware accelerators has powered the tremendous growth of machine learning in recent years. These accelerators deliver incomparable performance gains in processing high-volume matrix operators, particularly matrix…

Databases · Computer Science 2021-12-15 Yu-Ching Hu , Yuliang Li , Hung-Wei Tseng

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technologies include the Intel(TM) Xeon(TM) Data Center GPU Max…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-09 William E. Allcock , Benjamin S. Allen , James Anchell , Victor Anisimov , Thomas Applencourt , Abhishek Bagusetty , Ramesh Balakrishnan , Riccardo Balin , Solomon Bekele , Colleen Bertoni , Cyrus Blackworth , Renzo Bustamante , Kevin Canada , John Carrier , Christopher Chan-nui , Lance C. Cheney , Taylor Childers , Paul Coffman , Susan Coghlan , Tanima Dey , Michael D'Mello , Ashok Emani , Murali Emani , Kyle G. Felker , Sam Foreman , Olivier Franza , Longfei Gao , Marta García , María Garzarán , Balazs Gerofi , Yasaman Ghadar , Subrata Goswami , Neha Gupta , Kevin Harms , Väinö Hatanpää , Brian Holland , Carissa Holohan , Brian Homerding , Khalid Hossain , Xue Hu , Louise Huot , Huda Ibeid , Joseph A. Insley , Sai Jayanthi , Hong Jiang , Wei Jiang , Xiao-Yong Jin , Jeongnim Kim , Christopher Knight , Panagiotis Kourdis , Kalyan Kumaran , JaeHyuk Kwack , Janghaeng Lee , Ti Leggett , Ben Lenard , Chris Lewis , Nevin Liber , Johann Lombardi , Raymond M. Loy , Ye Luo , Bethany Lusch , Nilakantan Mahadevan , Beth Markey , Victor A. Mateevitsi , Gordon McPheeters , Ryan Milner , Jerome Mitchell , Vitali A. Morozov , Servesh Muralidharan , Tom Musta , Mrigendra Nagar , Vikram Narayana , Marieme Ngom , Anthony-Trung Nguyen , Nathan Nichols , Aditya Nishtala , James C. Osborn , Michael E. Papka , Scott Parker , Saumil S. Patel , Julia Piotrowska , Adrian C. Pope , Sucheta Raghunanda , Esteban Rangel , Paul M. Rich , Katherine M. Riley , Silvio Rizzi , Kris Rowe , Varuni Sastry , Adam Scovel , Filippo Simini , Haritha Siddabathuni Som , Patrick Steinbrecher , Rick Stevens , Xinmin Tian , Peter Upton , Thomas Uram , Archit K. Vasan , Álvaro Vázquez-Mayagoitia , Kaushik Velusamy , Brice Videau , Venkatram Vishwanath , Brian Whitney , Timothy J. Williams , Michael Woodacre , Sam Zeltner , Chuanjun Zhang , Gengbin Zheng , Huihuo Zheng

Modern scientific applications are getting more diverse, and the vector lengths in those applications vary widely. Contemporary Vector Processors (VPs) are designed either for short vector lengths, e.g., Fujitsu A64FX with 512-bit ARM SVE…

Spiking Neural Networks (SNNs) and transformers represent two powerful paradigms in neural computation, known for their low power consumption and ability to capture feature dependencies, respectively. However, transformer architectures…

Hardware Architecture · Computer Science 2025-03-27 Ching-Yao Chen , Meng-Chieh Chen , Tian-Sheuan Chang

Data movement is one of the main challenges of contemporary system architectures. Near-Data Processing (NDP) mitigates this issue by moving computation closer to the memory, avoiding excessive data movement. Our proposal, Vector-In-Memory…

Hardware Architecture · Computer Science 2022-03-29 Marco Antonio Zanata Alves , Sairo Santos , Aline S. Cordeiro , Francis B. Moreira , Paulo C. Santos , Luigi Carro

Tensor Cores have been an important unit to accelerate Fused Matrix Multiplication Accumulation (MMA) in all NVIDIA GPUs since Volta Architecture. To program Tensor Cores, users have to use either legacy wmma APIs or current mma APIs.…

Hardware Architecture · Computer Science 2022-11-29 Wei Sun , Ang Li , Tong Geng , Sander Stuijk , Henk Corporaal

The increasing diversity and complexity of transformer workloads at the edge present significant challenges in balancing performance, energy efficiency, and architectural flexibility. This paper introduces NX-CGRA, a programmable hardware…

Hardware Architecture · Computer Science 2025-11-24 Rohit Prasad

Real-time systems, particularly those used in domains like automated driving, are increasingly adopting neural networks. From this trend arises the need for high-performance hardware exhibiting predictable timing behavior. While…

Hardware Architecture · Computer Science 2026-02-26 Maximilian Kirschner , Konstantin Dudzik , Ben Krusekamp , Jürgen Becker

In recent years, interest in RISC-V computing architectures has moved from academic to mainstream, especially in the field of High Performance Computing where energy limitations are increasingly a concern. As of this year, the first single…

With the advancement of machine learning and deep learning, vector search becomes instrumental to many information retrieval systems, to search and find best matches to user queries based on their semantic similarities.These online services…

Computer Vision and Pattern Recognition · Computer Science 2018-09-13 Minjia Zhang , Yuxiong He

The objective of our research is to demonstrate the practical usage and orders of magnitude speedup of real-world applications by using alternative technologies to support high performance computing. Currently, the main barrier to the…

Astrophysics · Physics 2007-11-22 Robert J. Brunner , Volodymyr V. Kindratenko , Adam D. Myers
‹ Prev 1 2 3 10 Next ›