Performance · Computer Science
A model-driven approach for a new generation of adaptive libraries
Marco Cianfriglia, Flavio Vella, Cedric Nugteren, Anton Lokhmotov +1
2022-02-22
Distributed, Parallel, and Cluster Computing · Computer Science
Machine Learning Based Auto-tuning for Enhanced OpenCL Performance Portability
Thomas L. Falch, Anne C. Elster
2016-11-15
Distributed, Parallel, and Cluster Computing · Computer Science
Performance Optimization using Multimodal Modeling and Heterogeneous GNN
Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo Cesar +2
2023-04-28
Performance · Computer Science
Cross-Platform Performance Portability Using Highly Parametrized SYCL Kernels
John Lawson, Mehdi Goli, Duncan McBain, Daniel Soutar +1
2019-04-11
Performance · Computer Science
MLKAPS: Machine Learning and Adaptive Sampling for HPC Kernel Auto-tuning
Mathys Jam, Eric Petit, Pablo de Oliveira Castro, David Defour +2
2025-01-13
Distributed, Parallel, and Cluster Computing · Computer Science
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
L. A. Torres, Carlos J. Barrios H, Yves Denneulin
2024-05-28
Hardware Architecture · Computer Science
Heterogeneous Integration of In-Memory Analog Computing Architectures with Tensor Processing Units
Mohammed E. Elbtity, Brendan Reidy, Md Hasibul Amin, Ramtin Zand
2023-04-20
Distributed, Parallel, and Cluster Computing · Computer Science
A Benchmark Set of Highly-efficient CUDA and OpenCL Kernels and its Dynamic Autotuning with Kernel Tuning Toolkit
Filip Petrovič, David Střelák, Jana Hozzová, Jaroslav Oľha +3
2020-03-02
Quantum Physics · Physics
A Language and Hardware Independent Approach to Quantum-Classical Computing
Alexander J. McCaskey, Eugene F. Dumitrescu, Dmitry Liakh, Mengsu Chen +2
2018-08-01
Performance · Computer Science
COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPC
Ankur Lahiry, Banooqa Banday, Yugesh Bhattarai, Mohammad Zaeed +1
2026-04-28
Distributed, Parallel, and Cluster Computing · Computer Science
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
Baodi Shan, Mauricio Araya-Polo
2024-08-13
Distributed, Parallel, and Cluster Computing · Computer Science
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
Hang Liu, Junjie Li, Yinzhi Wang
2025-04-04
Distributed, Parallel, and Cluster Computing · Computer Science
Auto-Tuning High-Performance Programs Using Model Checking in Promela
Natalia Garanina, Sergey Staroletov, Sergei Gorlatch
2023-05-17
Distributed, Parallel, and Cluster Computing · Computer Science
AutoAccel: Automated Accelerator Generation and Optimization with Composable, Parallel and Pipeline Architecture
Jason Cong, Peng Wei, Cody Hao Yu, Peng Zhang
2018-09-24
Machine Learning · Computer Science
Better Trees: An empirical study on hyperparameter tuning of classification decision tree induction algorithms
Rafael Gomes Mantovani, Tomáš Horváth, André L. D. Rossi, Ricardo Cerri +3
2024-02-05
Distributed, Parallel, and Cluster Computing · Computer Science
Benchmarking optimization algorithms for auto-tuning GPU kernels
Richard Schoonhoven, Ben van Werkhoven, Kees Joost Batenburg
2022-10-05
Distributed, Parallel, and Cluster Computing · Computer Science
OpenCL Performance Prediction using Architecture-Independent Features
Beau Johnston, Greg Falzon, Josh Milthorpe
2018-11-04
Hardware Architecture · Computer Science
An In-Memory Analog Computing Co-Processor for Energy-Efficient CNN Inference on Mobile Devices
Mohammed Elbtity, Abhishek Singh, Brendan Reidy, Xiaochen Guo +1
2021-09-15
Distributed, Parallel, and Cluster Computing · Computer Science
Q-Learning Inspired Self-Tuning for Energy Efficiency in HPC
Andreas Gocht, Robert Schöne, Mario Bielert
2020-09-15
Performance · Computer Science
Crossing the Architectural Barrier: Evaluating Representative Regions of Parallel HPC Applications
Alexandra Ferreron, Radhika Jagtap, Sascha Bischoff, Roxana Rusitoru
2018-03-28
Distributed, Parallel, and Cluster Computing · Computer Science
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
Chendi Li, Haipeng Jia, Hang Cao, Jianyu Yao +5
2025-01-24