Distributed, Parallel, and Cluster Computing · Computer Science
To Use or Not to Use: CPUs' Cache Optimization Techniques on GPGPUs
Vajira Thambawita, Roshan G. Ragel, Dhammike Elkaduwe
2018-10-10
Distributed, Parallel, and Cluster Computing · Computer Science
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
Ehsan Yousefzadeh-Asl-Miandoab, Reza Karimzadeh, Danyal Yorulmaz, Bulat Ibragimov +1
2026-04-29
Distributed, Parallel, and Cluster Computing · Computer Science
Machine Learning Based Auto-tuning for Enhanced OpenCL Performance Portability
Thomas L. Falch, Anne C. Elster
2016-11-15
Machine Learning · Computer Science
A Study of Optimizations for Fine-tuning Large Language Models
Arjun Singh, Nikhil Pandey, Anup Shirgaonkar, Pavan Manoj +1
2024-06-07
Artificial Intelligence · Computer Science
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
Taeho Kim, Yanming Wang, Vatshank Chaturvedi, Lokesh Gupta +3
2024-04-18
Distributed, Parallel, and Cluster Computing · Computer Science
Going green: optimizing GPUs for energy efficiency through model-steered auto-tuning
Richard Schoonhoven, Bram Veenboer, Ben van Werkhoven, Kees Joost Batenburg
2022-11-15
Distributed, Parallel, and Cluster Computing · Computer Science
A Tool for Automatically Suggesting Source-Code Optimizations for Complex GPU Kernels
Saeed Taheri, Apan Qasem, Martin Burtscher
2019-10-18
Machine Learning · Computer Science
Towards making the most of NLP-based device mapping optimization for OpenCL kernels
Petros Vavaroutsos, Ioannis Oroutzoglou, Dimosthenis Masouros, Dimitrios Soudris
2022-08-31
Distributed, Parallel, and Cluster Computing · Computer Science
Benchmarking optimization algorithms for auto-tuning GPU kernels
Richard Schoonhoven, Ben van Werkhoven, Kees Joost Batenburg
2022-10-05
Performance · Computer Science
A Learned Performance Model for Tensor Processing Units
Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, Charith Mendis +3
2021-03-19
Hardware Architecture · Computer Science
Understanding Training Efficiency of Deep Learning Recommendation Models at Scale
Bilge Acun, Matthew Murphy, Xiaodong Wang, Jade Nie +2
2020-11-12
Hardware Architecture · Computer Science
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
Giuseppe M. Sarda, Nimish Shah, Debjyoti Bhattacharjee, Peter Debacker +1
2024-07-18
Distributed, Parallel, and Cluster Computing · Computer Science
A Benchmark Set of Highly-efficient CUDA and OpenCL Kernels and its Dynamic Autotuning with Kernel Tuning Toolkit
Filip Petrovič, David Střelák, Jana Hozzová, Jaroslav Oľha +3
2020-03-02
Distributed, Parallel, and Cluster Computing · Computer Science
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
Daniel Nichols, Konstantinos Parasyris, Charles Jekel, Abhinav Bhatele +1
2025-10-21
Distributed, Parallel, and Cluster Computing · Computer Science
Profiling and optimization of multi-card GPU machine learning jobs
Marcin Lawenda, Kyrylo Khloponin, Krzesimir Samborski, Łukasz Szustak
2025-05-30
Optimization and Control · Mathematics
A Framework for Self-Tuning Optimization Algorithm
Xin-She Yang, Suash Deb, M. Loomes, M. Karamanoglu
2013-12-20