Data Structures and Algorithms · Computer Science
Accelerating Graph Neural Networks with a Novel Matrix Compression Format
João N. F. Alves, Samir Moustafa, Siegfried Benkner, Alexandre P. Francisco +2
2024-09-05
Distributed, Parallel, and Cluster Computing · Computer Science
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
Mufakir Qamar Ansari, Mudabir Qamar Ansari
2025-07-30
Distributed, Parallel, and Cluster Computing · Computer Science
Efficient Strategies for Graph Pattern Mining Algorithms on GPUs
Samuel Ferraz, Vinicius Dias, Carlos H. C. Teixeira, George Teodoro +1
2022-12-12
Distributed, Parallel, and Cluster Computing · Computer Science
Reordering GPU Kernel Launches to Enable Efficient Concurrent Execution
Teng Li, Vikram K. Narayana, Tarek El-Ghazawi
2015-11-26
Distributed, Parallel, and Cluster Computing · Computer Science
Accelerating Maximal Biclique Enumeration on GPUs
Chou-Ying Hsieh, Chia-Ming Chang, Po-Hsiu Cheng, Sy-Yen Kuo
2025-05-23
Hardware Architecture · Computer Science
Towards a Multi-array Architecture for Accelerating Large-scale Matrix Multiplication on FPGAs
Junzhong Shen, Yuran Qiao, You Huang, Mei Wen +1
2018-03-13
Distributed, Parallel, and Cluster Computing · Computer Science
Optimizing CUDA Code By Kernel Fusion---Application on BLAS
J. Filipovič, M. Madzin, J. Fousek, L. Matyska
2017-09-13
Distributed, Parallel, and Cluster Computing · Computer Science
Multi-threaded Sparse Matrix-Matrix Multiplication for Many-Core and GPU Architectures
Mehmet Deveci, Christian Trott, Sivasankaran Rajamanickam
2018-01-10
Machine Learning · Computer Science
Kernel methods through the roof: handling billions of points efficiently
Giacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro Rudi
2020-11-30
Distributed, Parallel, and Cluster Computing · Computer Science
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
Haisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou +9
2025-01-17
Distributed, Parallel, and Cluster Computing · Computer Science
Faster and Cheaper: Parallelizing Large-Scale Matrix Factorization on GPUs
Wei Tan, Liangliang Cao, Liana Fong
2016-10-25
Machine Learning · Computer Science
Fast Polynomial Kernel Classification for Massive Data
Jinshan Zeng, Minrun Wu, Shao-Bo Lin, Ding-Xuan Zhou
2022-11-14
Machine Learning · Computer Science
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
Ruiqi Sun, Siwei Ye, Jie Zhao, Xin He +3
2024-08-21
Distributed, Parallel, and Cluster Computing · Computer Science
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
Zhonggen Li, Xiangyu Ke, Yifan Zhu, Yunjun Gao +1
2024-12-13
Machine Learning · Computer Science
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
Shaobo Ma, Chao Fang, Haikuo Shao, Zhongfeng Wang
2025-03-14
Computational Physics · Physics
Hybrid programming-model strategies for GPU offloading of electronic structure calculation kernels
Jean-Luc Fattebert, Christian F. A. Negre, Joshua Finkelstein, Jamaludin Mohd-Yusof +5
2024-01-26
Computational Physics · Physics
Performance Acceleration of Kernel Polynomial Method Applying Graphics Processing Units
Shixun Zhang, Shinichi Yamagiwa, Masahiko Okumura, Seiji Yunoki
2011-05-30
Distributed, Parallel, and Cluster Computing · Computer Science
Entropy Maximization in Sparse Matrix by Vector Multiplication ($\max_E SpMV$)
Paolo D'Alberto, Abhishek Jain, Ismail Bustany, Henri Fraisse +1
2023-08-02