English
Related papers

Related papers: Pinpoint resource allocation for GPU batch applica…

200 papers

In this paper, we evaluate training of deep recurrent neural networks with half-precision floats. We implement a distributed, data-parallel, synchronous training algorithm by integrating TensorFlow and CUDA-aware MPI to enable execution…

Machine Learning · Computer Science 2019-12-03 Alexey Svyatkovskiy , Julian Kates-Harbeck , William Tang

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for…

Instrumentation and Detectors · Physics 2026-03-12 Julia Gonski , Jenni Ott , Shiva Abbaszadeh , Sagar Addepalli , Matteo Cremonesi , Jennet Dickinson , Giuseppe Di Guglielmo , Erdem Yigit Ertorer , Lindsey Gray , Ryan Herbst , Christian Herwig , Tae Min Hong , Benedikt Maier , Maryam Bayat Makou , David Miller , Mark S. Neubauer , Cristián Peña , Dylan Rankin , Seon-Hee , Seo , Giordon Stark , Alexander Tapper , Audrey Corbeil Therrien , Ioannis Xiotidis , Keisuke Yoshihara , G Abarajithan , Sagar Addepalli , Nural Akchurin , Carlos Argüelles , Saptaparna Bhattacharya , Lorenzo Borella , Christian Boutan , Tom Braine , James Brau , Martin Breidenbach , Antonio Chahine , Talal Ahmed Chowdhury , Yuan-Tang Chou , Seokju Chung , Alberto Coppi , Mariarosaria D'Alfonso , Abhilasha Dave , Chance Desmet , Angela Di Fulvio , Karri DiPetrillo , Javier Duarte , Auralee Edelen , Jan Eysermans , Yongbin Feng , Emmett Forrestel , Dolores Garcia , Loredana Gastaldo , Julián García Pardiñas , Lino Gerlach , Loukas Gouskos , Katya Govorkova , Carl Grace , Christopher Grant , Philip Harris , Ciaran Hasnip , Timon Heim , Abraham Holtermann , Tae Min Hong , Gian Michele Innocenti , Koji Ishidoshiro , Miaochen Jin , Jyothisraj Johnson , Stephen Jones , Andreas Jung , Georgia Karagiorgi , Ryan Kastner , Nicholas Kamp , Doojin Kim , Kyoungchul Kong , Katie Kudela , Jelena Lalic , Bo-Cheng Lai , Yun-Tsung Lai , Tommy Lam , Jeffrey Lazar , Aobo Li , Zepeng Li , Haoyun Liu , Vladimir Lončar , Luca Macchiarulo , Christopher Madrid , Benedikt Maier , Zhenghua Ma , Prashansa Mukim , Mark S. Neubauer , Victoria Nguyen , Sungbin Oh , Isobel Ojalvo , Hideyoshi Ozaki , Simone Pagan Griso , Myeonghun Park , Christoph Paus , Santosh Parajuli , Benjamin Parpillon , Sara Pozzi , Ema Puljak , Benjamin Ramhorst , Amy Roberts , Larry Ruckman , Kate Scholberg , Sebastian Schmitt , Noah Singer , Eluned Anne Smith , Alexandre Sousa , Michael Spannowsky , Sioni Summers , Yanwen Sun , Daniel Tapia Takaki , Antonino Tumeo , Caterina Vernieri , Belina von Krosigk , Yash Vora , Linyan Wan , Michael H. L. S. Wang , Amanda Weinstein , Andy White , Simon Williams , Felix Yu

Modern LLM serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill and memory-bound decode phases. While current practices attempt to address this by organizing these phases…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-29 Zejia Lin , Hongxin Xu , Guanyi Chen , Zhiguang Chen , Yutong Lu , Xianwei Zhang

The surging demand for GPUs in datacenters for machine learning (ML) has made efficient GPU utilization crucial. However, meeting the diverse needs of ML models while optimizing resource usage is challenging. To enable transparent,…

With high-performance computing systems now running at exascale, optimizing power-scaling management and resource utilization has become more critical than ever. This paper explores runtime power-capping optimizations that leverage…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-26 Maria Patrou , Thomas Wang , Wael Elwasif , Markus Eisenbach , Ross Miller , William Godoy , Oscar Hernandez

Today's computing systems require moving data back-and-forth between computing resources (e.g., CPUs, GPUs, accelerators) and off-chip main memory so that computation can take place on the data. Unfortunately, this data movement is a major…

Hardware Architecture · Computer Science 2022-05-31 Geraldo F. Oliveira , Amirali Boroumand , Saugata Ghose , Juan Gómez-Luna , Onur Mutlu

Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different requests in the same batch to improve resource utilization…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-21 Amna Masood , Pratishtha Gaur , Nuwan Jayasena

This paper introduces a framework for solving alternating current optimal power flow (ACOPF) problems using graphics processing units (GPUs). While GPUs have demonstrated remarkable performance in various computing domains, their…

Optimization and Control · Mathematics 2026-05-11 Sungho Shin , François Pacaud , Mihai Anitescu

Recent years have seen the emergence of machine learning (ML) workloads deployed in warehouse-scale computing (WSC) settings, also known as ML fleets. As the computational demands placed on ML fleets have increased due to the rise of large…

One of the most important and commonly used operations in many linear algebra functions is matrix-matrix multiplication (GEMM), which is also a key component in obtaining high performance of many scientific codes. It is a computationally…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-18 Nenad Mijić , Davor Davidović

Recent years have seen the adoption of Machine Learning (ML) techniques to predict the performance of large-scale applications, mostly at a coarse level. In contrast, we propose to use ML techniques for performance prediction at a much…

In this paper we introduce the energy efficiency as a new metric for evaluating both hardware platforms based on Graphic Processor Units (GPU), and algorithm optimisations at High Energy Physics (HEP) experiments. We develop a method to…

High Energy Physics - Experiment · Physics 2026-05-01 Jiahui Zhuo , Arantza Oyanguren , Álvaro Fernández Casani , Luca Fiorini , Valerii Kholoimov

The widespread growth in LLM developments increasingly demands more computational power from clusters than what they can supply. Traditional LLM applications inherently require huge static resource allocations, which force users to either…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-17 Thanh Son Phung , Douglas Thain

Large-scale computing systems are increasingly using accelerators such as GPUs to enable peta- and exa-scale levels of compute to meet the needs of Machine Learning (ML) and scientific computing applications. Given the widespread and…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-20 Rutwik Jain , Brandon Tran , Keting Chen , Matthew D. Sinclair , Shivaram Venkataraman

The vision of super computer at every desk can be realized by powerful and highly parallel CPUs or GPUs or APUs. Graphics processors once specialized for the graphics applications only, are now used for the highly computational intensive…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-04-16 Chittampally Vasanth Raja , Srinivas Balasubramanian , Prakash S Raghavendra

Efficient scheduling of distributed deep learning (DL) jobs in large GPU clusters is crucial for resource efficiency and job performance. While server sharing among jobs improves resource utilization, interference among co-located DL jobs…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-12-28 Xiaoyang Zhao , Chuan Wu

In order to satisfy timing constraints, modern real-time applications require massively parallel accelerators such as General Purpose Graphic Processing Units (GPGPUs). Generation after generation, the number of computing clusters made…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-05-24 Houssam-Eddine Zahaf , Ignacio Sanudo Olmedo , Jayati Singh , Nicola Capodieci , Sebastien Faucou

Machine learning libraries such as TensorFlow and PyTorch simplify model implementation. However, researchers are still required to perform a non-trivial amount of manual tasks such as GPU allocation, training status tracking, and…

Modern GPU workloads increasingly demand efficient resource sharing, as many jobs do not require the full capacity of a GPU. Among sharing techniques, NVIDIA's Multi-Instance GPU (MIG) offers strong resource isolation by enabling…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-19 Hsu-Tzu Ting , Jerry Chou , Ming-Hung Chen , I-Hsin Chung

Highly parallelized workloads like machine learning training, inferences and general HPC tasks are greatly accelerated using GPU devices. In a cloud computing cluster, serving a GPU's computation power through multi-tasks sharing is highly…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-02-05 Wenqing Wu