Computational Physics · Physics
Performance of Kepler GTX Titan GPUs and Xeon Phi System
Hwancheol Jeong, Weonjong Lee, Jeonghwan Pak, Kwang-jong Choi +5
2013-11-05
High Energy Physics - Lattice · Physics
Code Optimization on Kepler GPUs and Xeon Phi
Yong-Chull Jang, Hwancheol Jeong, Jangho Kim, Weonjong Lee +2
2014-11-11
Distributed, Parallel, and Cluster Computing · Computer Science
Tiling for Performance Tuning on Different Models of GPUs
Chang Xu, Steven R. Kirk, Samantha Jenkins
2010-01-12
Distributed, Parallel, and Cluster Computing · Computer Science
Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking
Zhe Jia, Marco Maggioni, Benjamin Staiger, Daniele P. Scarpazza
2018-04-19
Hardware Architecture · Computer Science
Accelerator Codesign as Non-Linear Optimization
Nirmal Prajapati, Sanjay Rajopadhye, Hristo Djidjev, Nandkishore Santhi +2
2017-12-26
Distributed, Parallel, and Cluster Computing · Computer Science
Analyzing GPU Tensor Core Potential for Fast Reductions
Roberto Carrasco, Raimundo Vega, Cristóbal A. Navarro
2019-03-12
Distributed, Parallel, and Cluster Computing · Computer Science
Power and Energy-efficiency Roofline Model for GPUs
Millad Ghane, Jeff Larkin, Larry Shi, Sunita Chandrasekaran +1
2018-09-26
Computational Physics · Physics
Conjugate gradient solvers on Intel Xeon Phi and NVIDIA GPUs
O. Kaczmarek, C. Schmidt, P. Steinbrecher, M. Wagner
2014-11-18
Distributed, Parallel, and Cluster Computing · Computer Science
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
Aaron Jarmusch, Nathan Graddon, Sunita Chandrasekaran
2025-07-23
Computer Vision and Pattern Recognition · Computer Science
GPGPU Acceleration of the KAZE Image Feature Extraction Algorithm
Ramkumar B, R. S. Hegde, Rob Laber, Hristo Bojinov
2017-06-22
Distributed, Parallel, and Cluster Computing · Computer Science
BitGNN: Unleashing the Performance Potential of Binary Graph Neural Networks on GPUs
Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Sutanay Choudhury +1
2023-06-06
Distributed, Parallel, and Cluster Computing · Computer Science
NVIDIA Tensor Core Programmability, Performance & Precision
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng +1
2018-12-18
Distributed, Parallel, and Cluster Computing · Computer Science
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du +3
2025-09-05
Distributed, Parallel, and Cluster Computing · Computer Science
A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
Yixing Li, Zichuan Liu, Kai Xu, Hao Yu +1
2017-06-09
Distributed, Parallel, and Cluster Computing · Computer Science
Kernelet: High-Throughput GPU Kernel Executions with Dynamic Slicing and Scheduling
Jianlong Zhong, Bingsheng He
2013-03-22
Hardware Architecture · Computer Science
Performance Analysis and Optimization Opportunities for NVIDIA Automotive GPUs
Hamid Tabani, Fabio Mazzocchetti, Pedro Benedicte, Jaume Abella +1
2021-04-19
Hardware Architecture · Computer Science
Exploring FPGA designs for MX and beyond
Ebby Samson, Naveen Mellempudi, Wayne Luk, George A. Constantinides
2024-07-02
Hardware Architecture · Computer Science
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du +2
2024-02-22