Distributed, Parallel, and Cluster Computing · Computer Science
Performance-Portable Many-Core Plasma Simulations: Porting PIConGPU to OpenPower and Beyond
Erik Zenker, René Widera, Axel Huebl, Guido Juckeland +3
2016-11-07
Distributed, Parallel, and Cluster Computing · Computer Science
A Review of CUDA, MapReduce, and Pthreads Parallel Computing Models
Kato Mivule, Benjamin Harvey, Crystal Cobb, Hoda El Sayed
2014-10-17
Accelerator Physics · Physics
Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels
Ka Hei Martin Kwok, Matti Kortelainen, Giuseppe Cerati, Alexei Strelchenko +10
2024-01-26
Distributed, Parallel, and Cluster Computing · Computer Science
Taking GPU Programming Models to Task for Performance Portability
Joshua H. Davis, Pranav Sivaraman, Joy Kitson, Konstantinos Parasyris +4
2025-09-08
Distributed, Parallel, and Cluster Computing · Computer Science
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
Baodi Shan, Mauricio Araya-Polo
2024-08-13
Distributed, Parallel, and Cluster Computing · Computer Science
Parallel Programming Models for Heterogeneous Many-Cores : A Survey
Jianbin Fang, Chun Huang, Tao Tang, Zheng Wang
2020-05-11
Distributed, Parallel, and Cluster Computing · Computer Science
Portability for GPU-accelerated molecular docking applications for cloud and HPC: can portable compiler directives provide performance across all platforms?
Mathialakan Thavappiragasam, Wael Elwasif, Ada Sedova
2022-03-07
Performance · Computer Science
Cross-Platform Performance Portability Using Highly Parametrized SYCL Kernels
John Lawson, Mehdi Goli, Duncan McBain, Daniel Soutar +1
2019-04-11
Distributed, Parallel, and Cluster Computing · Computer Science
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
Wael Elwasif, William Godoy, Nick Hagerty, J. Austin Harris +27
2024-09-20
Distributed, Parallel, and Cluster Computing · Computer Science
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
Nizar ALHafez, Ahmad Kurdi
2025-06-19
High Energy Physics - Experiment · Physics
Exploring code portability solutions for HEP with a particle tracking test code
Hammad Ather, Sophie Berkman, Giuseppe Cerati, Matti Kortelainen +9
2024-09-17
Distributed, Parallel, and Cluster Computing · Computer Science
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
Manuel Costanzo, Enzo Rucci, Carlos García-Sánchez, Marcelo Naiouf +1
2025-04-15
Distributed, Parallel, and Cluster Computing · Computer Science
Execution of Compound Multi-Kernel OpenCL Computations in Multi-CPU/Multi-GPU Environments
Fábio Soldado, Fernando Alexandre, Hervé Paulino
2015-10-23
Distributed, Parallel, and Cluster Computing · Computer Science
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
Panagiotis-Eleftherios Eleftherakis, George Anagnostopoulos, Anastassis Kapetanakis, Mohammad Umair +9
2026-01-21
Distributed, Parallel, and Cluster Computing · Computer Science
CPU and/or GPU: Revisiting the GPU Vs. CPU Myth
Kishore Kothapalli, Dip Sankar Banerjee, P. J. Narayanan, Surinder Sood +12
2013-03-12
Distributed, Parallel, and Cluster Computing · Computer Science
Performance and Portability of Accelerated Lattice Boltzmann Applications with OpenACC
E. Calore, A. Gabbana, J. Kraus, S. F. Schifano +1
2017-03-02
Distributed, Parallel, and Cluster Computing · Computer Science
GPGPU Processing in CUDA Architecture
Jayshree Ghorpade, Jitendra Parande, Madhura Kulkarni, Amit Bawaskar
2012-02-21
Distributed, Parallel, and Cluster Computing · Computer Science
Runtime Support for Performance Portability on Heterogeneous Distributed Platforms
Polykarpos Thomadakis, Nikos Chrisochoides
2023-03-09
Programming Languages · Computer Science
High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel Constructs
William S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo +2
2022-07-04
Computational Physics · Physics
Efficient molecular dynamics simulations with many-body potentials on graphics processing units
Zheyong Fan, Wei Chen, Ville Vierimaa, Ari Harju
2017-06-27
Hardware Architecture · Computer Science
Performance monitoring for multicore embedded computing systems on FPGAs
Lesley Shannon, Eric Matthews, Nicholas Doyle, Alexandra Fedorova
2015-08-31
Distributed, Parallel, and Cluster Computing · Computer Science
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
Mufakir Qamar Ansari, Mudabir Qamar Ansari
2025-07-30