English
Related papers

Related papers: Tarema: Adaptive Resource Allocation for Scalable …

200 papers

Problem Definition: Allocating sufficient capacity to cloud services is a challenging task, especially when demand is time-varying, heterogeneous, contains batches, and requires multiple types of resources for processing. In this setting,…

Applications · Statistics 2022-09-21 Eugene Furman , Arik Senderovich , Shane Bergsma , J. Christopher Beck

Quantum computing resources are increasingly being incorporated into high-performance computing (HPC) environments as co-processors for hybrid workloads. To support this paradigm, quantum devices must be treated as schedulable first-class…

Data center schedulers operate at unprecedented scales today to accommodate the growing demand for computing and storage power. The challenge that schedulers face is meeting the requirements of scheduling speeds despite the scale. To do so,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-03-05 Meghana Thiyyakat , Subramaniam Kalambur , Rishit Chaudhary , Saurav G Nayak , Adarsh Shetty , Dinkar Sitaram

A heterogeneous memory has a single address space with fast access to some addresses (a fast tier of DRAM) and slow access to other addresses (a capacity tier of CXL-attached memory or NVM). A tiered memory system aims to maximize the…

Emerging Technologies · Computer Science 2025-10-28 Rohan Kadekodi , Haoran Peng , Gilbert Bernstein , Michael D. Ernst , Baris Kasikci

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from execution of a…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-04-04 Rafael Ferreira da Silva , Rosa M. Badia , Venkat Bala , Debbie Bard , Peer-Timo Bremer , Ian Buckley , Silvina Caino-Lores , Kyle Chard , Carole Goble , Shantenu Jha , Daniel S. Katz , Daniel Laney , Manish Parashar , Frederic Suter , Nick Tyler , Thomas Uram , Ilkay Altintas , Stefan Andersson , William Arndt , Juan Aznar , Jonathan Bader , Bartosz Balis , Chris Blanton , Kelly Rosa Braghetto , Aharon Brodutch , Paul Brunk , Henri Casanova , Alba Cervera Lierta , Justin Chigu , Taina Coleman , Nick Collier , Iacopo Colonnelli , Frederik Coppens , Michael Crusoe , Will Cunningham , Bruno de Paula Kinoshita , Paolo Di Tommaso , Charles Doutriaux , Matthew Downton , Wael Elwasif , Bjoern Enders , Chris Erdmann , Thomas Fahringer , Ludmilla Figueiredo , Rosa Filgueira , Martin Foltin , Anne Fouilloux , Luiz Gadelha , Andy Gallo , Artur Garcia Saez , Daniel Garijo , Roman Gerlach , Ryan Grant , Samuel Grayson , Patricia Grubel , Johan Gustafsson , Valerie Hayot-Sasson , Oscar Hernandez , Marcus Hilbrich , AnnMary Justine , Ian Laflotte , Fabian Lehmann , Andre Luckow , Jakob Luettgau , Ketan Maheshwari , Motohiko Matsuda , Doriana Medic , Pete Mendygral , Marek Michalewicz , Jorji Nonaka , Maciej Pawlik , Loic Pottier , Line Pouchard , Mathias Putz , Santosh Kumar Radha , Lavanya Ramakrishnan , Sashko Ristov , Paul Romano , Daniel Rosendo , Martin Ruefenacht , Katarzyna Rycerz , Nishant Saurabh , Volodymyr Savchenko , Martin Schulz , Christine Simpson , Raul Sirvent , Tyler Skluzacek , Stian Soiland-Reyes , Renan Souza , Sreenivas Rangan Sukumar , Ziheng Sun , Alan Sussman , Douglas Thain , Mikhail Titov , Benjamin Tovar , Aalap Tripathy , Matteo Turilli , Bartosz Tuznik , Hubertus van Dam , Aurelio Vivas , Logan Ward , Patrick Widener , Sean Wilkinson , Justyna Zawalska , Mahnoor Zulfiqar

Analyzing large datasets with distributed dataflow systems requires the use of clusters. Public cloud providers offer a large variety and quantity of resources that can be used for such clusters. However, picking the appropriate resources…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-04-28 Jonathan Will , Jonathan Bader , Lauritz Thamsen

Cloud providers usually offer diverse types of hardware for their users. Customers exploit this option to deploy cloud instances featuring GPUs, FPGAs, architectures other than x86 (e.g., ARM, IBM Power8), or featuring certain specific…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-06-28 Isabelly Rocha , Christian Göttel , Pascal Felber , Marcelo Pasin , Romain Rouvoy , Valerio Schiavoni

Edge computing enables latency-critical applications to process data close to end devices, yet task heterogeneity and limited resources pose significant challenges to efficient orchestration. This paper presents a measurement-driven,…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-01 Yongmin Zhang , Pengyu Huang , Mingyi Dong , Jing Yao

Molecular dynamics (MD) simulations are widely used to study large-scale molecular systems. HPC systems are ideal platforms to run these studies, however, reaching the necessary simulation timescale to detect rare processes is challenging,…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-08-22 Tu Mai Anh Do , Loïc Pottier , Rafael Ferreira da Silva , Frédéric Suter , Silvina Caíno-Lores , Michela Taufer , Ewa Deelman

We consider a parallel system of $m$ identical machines prone to unpredictable crashes and restarts, trying to cope with the continuous arrival of tasks to be executed. Tasks have different computational requirements (i.e., processing time…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-03-21 Elli Zavou , Antonio Fernández Anta

Asynchronous methods are fundamental for parallelizing computations in distributed machine learning. They aim to accelerate training by fully utilizing all available resources. However, their greedy approach can lead to inefficiencies using…

Machine Learning · Computer Science 2025-05-23 Artavazd Maranjyan , El Mehdi Saad , Peter Richtárik , Francesco Orabona

Low-Rank Adaptation (LoRA) is widely adopted for downstream fine-tuning of foundation models due to its efficiency and zero additional inference cost. Many real-world applications require foundation models to specialize in several specific…

Machine Learning · Computer Science 2025-09-30 Jian Liang , Wenke Huang , Xianda Guo , Guancheng Wan , Bo Du , Mang Ye

Increasing data volumes in scientific experiments necessitate the use of high-performance computing (HPC) resources for data analysis. In many scientific fields, the data generated from scientific instruments and supercomputer simulations…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-25 Sam Nickolay , Eun-Sung Jung , Rajkumar Kettimuthu , Ian Foster

Human involvement is critical in training and deploying AI systems in high-stakes defence and security contexts. However, real-time interaction is impractical in HPC environments due to compute intensity and resource constraints. We present…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-06 Sergio Mendoza , Cedric Bhihe , Natalia Zamora , David Modesto , Jose Martin Bugallo Batalla , Jesus Gomez Canovas , Rafel Palomo Avellaneda , Miguel Perez Espinosa

Adaptive workloads can change on--the--fly the configuration of their jobs, in terms of number of processes. In order to carry out these job reconfigurations, we have designed a methodology which enables a job to communicate with the…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-06-01 Sergio Iserte , Rafael Mayo , Enrique S. Quintana-Orti , Vicenc Beltran , Antonio J. Peña

The growing complexity and scale of scientific workflows in high performance computing (HPC) environments have led to significant challenges in managing energy consumption without compromising computational performance. Traditional…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-25 Ali Zahir , Ashiq Anjum , Mark Wilkinson , Jeyan Thiyagalingam

Training large language models (LLMs) in the cloud faces growing memory bottlenecks due to the limited capacity and high cost of GPUs. While GPU memory offloading to CPU and NVMe has made large-scale training more feasible, existing…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-19 Sabiha Afroz , Redwan Ibne Seraj Khan , Hadeel Albahar , Jingoo Han , Ali R. Butt

Clustering algorithms are iterative and have complex data access patterns that result in many small random memory accesses. The performance of parallel implementations suffer from synchronous barriers for each iteration and skewed…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-01-19 Disa Mhembere , Da Zheng , Carey E. Priebe , Joshua T. Vogelstein , Randal Burns

The rapid growth of large language model (LLM) services imposes increasing demands on distributed GPU inference infrastructure. Most existing scheduling systems follow a reactive paradigm, relying solely on the current system state to make…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-17 Chengze Du , Zhiwei Yu , Heng Xu , Haojie Wang , Bo liu , Jialong Li

Transformer-based models are becoming deeper and larger recently. For better scalability, an underlying training solution in industry is to split billions of parameters (tensors) into many tasks and then run them across homogeneous…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-01-23 Zhigang Wang , Xu Zhang , Ning Wang , Chuanfei Xu , Jie Nie , Zhiqiang Wei , Yu Gu , Ge Yu