English
Related papers

Related papers: Preparing Ginkgo for AMD GPUs -- A Testimonial on …

200 papers

Processing-In-Memory (PIM) architectures offer a promising approach to accelerate Graph Neural Network (GNN) training and inference. However, various PIM devices such as ReRAM, FeFET, PCM, MRAM, and SRAM exist, with each device offering…

GPUs have become the dominant source of computing power for high performance computing and are increasingly being used across the High Energy Physics computing landscape for a wide variety of tasks. Though NVIDIA is currently the main…

As part of the Exascale Computing Project (ECP), a recent focus of development efforts for the SUite of Nonlinear and DIfferential/ALgebraic equation Solvers (SUNDIALS) has been to enable GPU-accelerated time integration in scientific…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-04 Cody J. Balos , David J. Gardner , Carol S. Woodward , Daniel R. Reynolds

Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to dependencies on both GPU kernel efficiency and host-side settings. Although LLM-based methods show…

Multiagent Systems · Computer Science 2026-03-04 Shiyang Li , Zijian Zhang , Winson Chen , Yuebo Luo , Mingyi Hong , Caiwen Ding

Mixed-Integer Programming (MIP), particularly Mixed-Integer Linear Programming (MILP) and Mixed-Integer Quadratic Programming (MIQP), has found extensive applications in domains such as portfolio optimization and network flow control, which…

Optimization and Control · Mathematics 2026-02-03 Zayn Wang

Quantization is a key technique to reduce the resource requirement and improve the performance of neural network deployment. However, different hardware backends such as x86 CPU, NVIDIA GPU, ARM CPU, and accelerators may demand different…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Ziheng Jiang , Animesh Jain , Andrew Liu , Josh Fromm , Chengqian Ma , Tianqi Chen , Luis Ceze

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper's memory subsystem, highlighting improvements in the L2…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-05 Weile Luo , Ruibo Fan , Zeyu Li , Dayou Du , Hongyuan Liu , Qiang Wang , Xiaowen Chu

This paper presents the implementation of a HLLC finite volume solver using GPU technology for the solution of shallow water problems in two dimensions. It compares both CPU and GPU approaches for implementing all the solver's steps. The…

Computational Engineering, Finance, and Science · Computer Science 2018-07-03 Fabrice Zaoui

Modern Graphics Processing Units (GPUs) are well provisioned to support the concurrent execution of thousands of threads. Unfortunately, different bottlenecks during execution and heterogeneous application requirements create imbalances in…

To assess how future progress in gravitational microlensing computation at high optical depth will rely on both hardware and software solutions, we compare a direct inverse ray-shooting code implemented on a graphics processing unit (GPU)…

Instrumentation and Methods for Astrophysics · Physics 2015-05-19 N. F. Bate , C. J. Fluke , B. R. Barsdell , H. Garsden , G. F. Lewis

An existing hybrid MPI-OpenMP scheme is augmented with a CUDA-based fine grain parallelization approach for multidimensional distributed Fourier transforms, in a well-characterized pseudospectral fluid turbulence code. Basics of the hybrid…

Computational Physics · Physics 2018-08-07 Duane Rosenberg , Pablo D. Mininni , Raghu Reddy , Annick Pouquet

Geometric Semantic Genetic Programming (GSGP) is a state-of-the-art machine learning method based on evolutionary computation. GSGP performs search operations directly at the level of program semantics, which can be done more efficiently…

Neural and Evolutionary Computing · Computer Science 2021-06-09 Leonardo Trujillo , Jose Manuel Muñoz Contreras , Daniel E Hernandez , Mauro Castelli , Juan J Tapia

GPUs are playing an increasingly important role in general-purpose computing. Many algorithms require synchronizations at different levels of granularity in a single GPU. Additionally, the emergence of dense GPU nodes also calls for…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-04-14 Lingqi Zhang , Mohamed Wahib , Haoyu Zhang , Satoshi Matsuoka

The field of plasma physics heavily relies on simulations to model various phenomena, such as instabilities, turbulence, and nonlinear behaviors that would otherwise be difficult to study from a purely theoretical approach. Simulations are…

Plasma Physics · Physics 2026-02-05 Giorgio Daneri

Modern GPUs incorporate specialized matrix units such as Tensor Cores to accelerate GEMM operations, which are central to deep learning workloads. However, existing matrix unit designs are tightly coupled to the SIMT core, restricting…

Hardware Architecture · Computer Science 2025-03-04 Hansung Kim , Ruohan Richard Yan , Joshua You , Tieliang Vamber Yang , Yakun Sophia Shao

To deploy large Mixture-of-Experts (MoE) models cost-effectively, offloading-based single-GPU heterogeneous inference is crucial. While GPU-CPU architectures that offload cold experts are constrained by host memory bandwidth, emerging…

Hardware Architecture · Computer Science 2026-03-03 Yudong Pan , Yintao He , Tianhua Han , Lian Liu , Shixin Zhao , Zhirong Chen , Mengdi Wang , Cangyuan Li , Yinhe Han , Ying Wang

Modern GPU systems are constantly evolving to meet the needs of computing-intensive applications in scientific and machine learning domains. However, there is typically a gap between the hardware capacity and the achievable application…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-02 Gabin Schieffer , Ruimin Shi , Stefano Markidis , Andreas Herten , Jennifer Faj , Ivy Peng

We present here a set of examples, classes and tools which can be used for statistical analysis in Graphics Processing Units (GPU). This includes binned and unbinned maximum likelihood fits, pseudo-experiment generation, convolutions,…

The LHCb experiment at CERN is undergoing an upgrade in preparation for the Run 3 data taking period of the LHC. As part of this upgrade the trigger is moving to a fully software implementation operating at the LHC bunch crossing rate. We…

Instrumentation and Detectors · Physics 2022-01-06 R. Aaij , M. Adinolfi , S. Aiola , S. Akar , J. Albrecht , M. Alexander , S. Amato , Y. Amhis , F. Archilli , M. Bala , G. Bassi , L. Bian , M. P. Blago , T. Boettcher , A. Boldyrev , S. Borghi , A. Brea Rodriguez , L. Calefice , M. Calvo Gomez , D. H. Cámpora Pérez , A. Cardini , M. Cattaneo , V. Chobanova , G. Ciezarek , X. Cid Vidal , J. L. Cobbledick , J. A. B. Coelho , T. Colombo , A. Contu , B. Couturier , D. C. Craik , R. Currie , P. d'Argent , M. De Cian , D. Derkach , F. Dordei , M. Dorigo , L. Dufour , P. Durante , A. Dziurda , A. Dzyuba , S. Easo , S. Esen , P. Fernandez Declara , S. Filippov , C. Fitzpatrick , M. Frank , P. Gandini , V. V. Gligorov , E. Golobardes , G. Graziani , L. Grillo , P. A. Günther , S. Hansmann-Menzemer , A. M. Hennequin , L. Henry , D. Hill , S. E. Hollitt , J. Hu , W. Hulsbergen , R. J. Hunter , M. Hushchyn , B. K. Jashal , C. R. Jones , S. Klaver , K. Klimaszewski , R. Kopecna , W. Krzemien , M. Kucharczyk , R. Lane , F. Lazzari , R. Le Gac , P. Li , J. H. Lopes , M. Lucio Martinez , A. Lupato , O. Lupton , X. Lyu , F. Machefert , O. Madejczyk , S. Malde , J. F. Marchand , S. Mariani , C. Marin Benito , D. Martinez Santos , F. Martinez Vidal , R. Matev , M. Mazurek , B. Mitreska , D. S. Mitzel , M. J. Morello , H. Mu , P. Muzzetto , P. Naik , M. Needham , N. Neri , N. Neufeld , N. S. Nolte , D. O'Hanlon , A. Oyanguren , M. Pepe Altarelli , S. Petrucci , M. Petruzzo , L. Pica , F. Pisani , A. Piucci , F. Polci , A. Poluektov , E. Polycarpo , C. Prouve , G. Punzi , R. Quagliani , R. I. Rabadan Trejo , M. Ramos Pernas , M. S. Rangel , F. Ratnikov , G. Raven , F. Reiss , V. Renaudin , P. Robbe , A. Ryzhikov , M. Santimaria , M. Saur , M. Schiller , R. Schwemmer , B. Sciascia , A. Solomin , F. Suljik , N. Skidmore , M. D. Sokoloff , P. Spradlin , M. Stahl , S. Stahl , H. Stevens , L. Sun , A. Szabelski , T. Szumlak , M. Szymanski , D. Y. Tou , G. Tuci , A. Usachov , N. Valls Canudas , R. Vazquez Gomez , S. Vecchi , M. Vesterinen , X. Vilasis-Cardona , D. Vom Bruch , Z. Wang , T. Wojton , M. Whitehead , M. Williams , M. Witek , Y. Xie , A. Xu , H. Yin , M. Zdybal , O. Zenaiev , D. Zhang , L. Zhang , X. Zhu

Effective Hamiltonian calculations for large quantum systems can be both analytically intractable and numerically expensive using standard techniques. In this manuscript, we present numerical techniques inspired by Nonperturbative…