English
Related papers

Related papers: Adventures with Grace Hopper AI Super Chip and the…

200 papers

The National Research Platform (NRP) represents a distributed, multi-tenant Kubernetes-based cyberinfrastructure designed to facilitate collaborative scientific computing. Spanning over 75 locations in the U.S. and internationally, the NRP…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-30 Derek Weitzel , Ashton Graves , Sam Albin , Huijun Zhu , Frank Würthwein , Mahidhar Tatineni , Dmitry Mishin , John Graham , Elham E Khoda , Mohammad Firas Sada , Larry Smarr , Thomas DeFanti

Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-27 Luigi Fusco , Mikhail Khalilov , Marcin Chrapek , Giridhar Chukkapalli , Thomas Schulthess , Torsten Hoefler

The objective of our research is to demonstrate the practical usage and orders of magnitude speedup of real-world applications by using alternative technologies to support high performance computing. Currently, the main barrier to the…

Astrophysics · Physics 2007-11-22 Robert J. Brunner , Volodymyr V. Kindratenko , Adam D. Myers

Significant investments to upgrade and construct large-scale scientific facilities demand commensurate investments in R&D to design algorithms and computing approaches to enable scientific and engineering breakthroughs in the big data era.…

The rapid growth of AI, data-intensive science, and digital twin technologies has driven an unprecedented demand for high-performance computing (HPC) across the research ecosystem. While national laboratories and industrial hyperscalers…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-25 Peng Shu , Junhao Chen , Zhengliang Liu , Huaqin Zhao , Xinliang Li , Tianming Liu

Graphics processing units (GPUs) are continually evolving to cater to the computational demands of contemporary general-purpose workloads, particularly those driven by artificial intelligence (AI) utilizing deep learning techniques. A…

Hardware Architecture · Computer Science 2024-02-22 Weile Luo , Ruibo Fan , Zeyu Li , Dayou Du , Qiang Wang , Xiaowen Chu

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper's memory subsystem, highlighting improvements in the L2…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-05 Weile Luo , Ruibo Fan , Zeyu Li , Dayou Du , Hongyuan Liu , Qiang Wang , Xiaowen Chu

This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-30 Mohammad Firas Sada , John J. Graham , Elham E Khoda , Mahidhar Tatineni , Dmitry Mishin , Rajesh K. Gupta , Rick Wagner , Larry Smarr , Thomas A. DeFanti , Frank Würthwein

A range of computational biology software (GROMACS, AMBER, NAMD, LAMMPS, OpenMM, Psi4 and RELION) was benchmarked on a representative selection of HPC hardware, including AMD EPYC 7742 CPU nodes, NVIDIA V100 and AMD MI250X GPU nodes, and an…

Memory management across discrete CPU and GPU physical memory is traditionally achieved through explicit GPU allocations and data copy or unified virtual memory. The Grace Hopper Superchip, for the first time, supports an integrated CPU-GPU…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-11 Gabin Schieffer , Jacob Wahlgren , Jie Ren , Jennifer Faj , Ivy Peng

Field Programmable Gate Arrays (FPGAs) plays an increasingly important role in data sampling and processing industries due to its highly parallel architecture, low power consumption, and flexibility in custom algorithms. Especially, in the…

Computer Vision and Pattern Recognition · Computer Science 2017-11-17 Yufeng Hao

FPGA is appropriate for fix-point neural networks computing due to high power efficiency and configurability. However, its design must be intensively refined to achieve high performance using limited hardware resources. We present an…

Hardware Architecture · Computer Science 2022-01-03 Qingyang Yi , Heming Sun , Masahiro Fujita

The practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present…

Hardware Architecture · Computer Science 2022-09-27 Chaolin Rao , Huangjie Yu , Haochuan Wan , Jindong Zhou , Yueyang Zheng , Yu Ma , Anpei Chen , Minye Wu , Binzhe Yuan , Pingqiang Zhou , Xin Lou , Jingyi Yu

The increasing interest in the usage of Artificial Intelligence techniques (AI) from the research community and industry to tackle "real world" problems, requires High Performance Computing (HPC) resources to efficiently compute and scale…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-14 David Brayford , Sofia Vallecorsa , Atanas Atanasov , Fabio Baruffa , Walter Riviera

Artificial intelligence (AI) research today is largely driven by ever-larger neural network models trained on graphics processing units (GPUs). This paradigm has yielded remarkable progress, but it also risks entrenching a hardware lottery…

Artificial Intelligence · Computer Science 2025-11-17 Bipin Rajendran , Osvaldo Simeone , Bashir M. Al-Hashimi

Many architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU)---deployed in datacenters since 2015 that…

We present both a novel Convolutional Neural Network (CNN) accelerator architecture and a network compiler for FPGAs that outperforms all prior work. Instead of having generic processing elements that together process one layer at a time,…

Hardware Architecture · Computer Science 2020-07-22 Mathew Hall , Vaughn Betz

The pursuit of many research questions requires massive computational resources. State-of-the-art research in physical processes using simulations, the training of neural networks for deep learning, or the analysis of big data are all…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-08 Magnus Själander , Magnus Jahre , Gunnar Tufte , Nico Reissmann

Applications like Big Data, Machine Learning, Deep Learning and even other Engineering and Scientific research requires a lot of computing power; making High-Performance Computing (HPC) an important field. But access to Supercomputers is…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-03-03 Hamza Ali Imran , Saad Wazir , Ahmed Jamal Ikram , Ataul Aziz Ikram , Hanif Ullah , Maryam Ehsan

The IBM Neural Computer (INC) is a highly flexible, re-configurable parallel processing system that is intended as a research and development platform for emerging machine intelligence algorithms and computational neuroscience. It consists…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-03-26 Pritish Narayanan , Charles E. Cox , Alexis Asseman , Nicolas Antoine , Harald Huels , Winfried W. Wilcke , Ahmet S. Ozcan
‹ Prev 1 2 3 10 Next ›