English
Related papers

Related papers: Tarema: Adaptive Resource Allocation for Scalable …

200 papers

This paper investigates co-scheduling algorithms for processing a set of parallel applications. Instead of executing each application one by one, using a maximum degree of parallelism for each of them, we aim at scheduling several…

Data Structures and Algorithms · Computer Science 2013-05-01 Guillaume Aupy , Manu Shantharam , Anne Benoit , Yves Robert , Padma Raghavan

Modern distributed machine learning (ML) training workloads benefit significantly from leveraging GPUs. However, significant contention ensues when multiple such workloads are run atop a shared cluster of GPUs. A key question is how to…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-30 Kshiteej Mahajan , Arjun Balasubramanian , Arjun Singhvi , Shivaram Venkataraman , Aditya Akella , Amar Phanishayee , Shuchi Chawla

As Kubernetes becomes the infrastructure of the cloud-native era, the integration of workflow systems with Kubernetes is gaining more and more popularity. To our knowledge, workflow systems employ scheduling algorithms that optimize task…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-07-07 Chenggang Shan , Guan Wang , Yuanqing Xia , Yufeng Zhan , Jinhui Zhang

Many organizations routinely analyze large datasets using systems for distributed data-parallel processing and clusters of commodity resources. Yet, users need to configure adequate resources for their data processing jobs. This requires…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-06-02 Lauritz Thamsen , Dominik Scheinert , Jonathan Will , Jonathan Bader , Odej Kao

Task-based programming models have become very popular, as they offer an attractive solution to parallelize serial application code with task and data annotations. They usually depend on a runtime system that schedules the tasks to multiple…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-06-15 Spyros Lyberis , Polyvios Pratikakis , Iakovos Mavroidis , Dimitrios S. Nikolopoulos

Database platform-as-a-service (dbPaaS) is developing rapidly and a large number of databases have been migrated to run on the Clouds for the low cost and flexibility. Emerging Clouds rely on the tenants to provide the resource…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-12-30 Ningxin Zheng , Quan Chen , Yong Yang , Wei Zhang , Jin Li , Wenli Zheng , Minyi Guo

Task parallelism is designed to simplify the task of parallel programming. When executing a task parallel program on modern NUMA architectures, it can fail to scale due to the phenomenon called work inflation, where the overall processing…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-01-08 Justin Deters , Jiaye Wu , Yifan Xu , I-Ting Angelina Lee

Test-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current methods employ uniform…

Computation and Language · Computer Science 2025-12-02 Yang Xiao , Chunpu Xu , Ruifeng Yuan , Jiashuo Wang , Wenjie Li , Pengfei Liu

In this study, a cluster-computing environment is employed as a computational platform. In order to increase the efficiency of the system, a dynamic task scheduling algorithm is proposed, which balances the load among the nodes of the…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-02-22 I. K. Savvas , M. Tahar Kechadi

The ever increasing adoption of mobile devices with limited energy storage capacity, on the one hand, and more awareness of the environmental impact of massive data centres and server pools, on the other hand, have both led to an increased…

Discrete Mathematics · Computer Science 2018-06-14 Rodrigo A. Carrasco , Garud Iyengar , Cliff Stein

As large scale cloud computing centers become more popular than individual servers, predicting future resource demand need has become an important problem. Forecasting resource need allows public cloud providers to proactively allocate or…

Machine Learning · Computer Science 2020-07-17 Langston Nashold , Rayan Krishnan

The emerging edge-hub-cloud paradigm has enabled the development of innovative latency-critical cyber-physical applications in the edge-cloud continuum. However, this paradigm poses multiple challenges due to the heterogeneity of the…

Networking and Internet Architecture · Computer Science 2025-11-18 Andreas Kouloumpris , Georgios L. Stavrinides , Maria K. Michael , Theocharis Theocharides

Heterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is…

Cryptography and Security · Computer Science 2024-07-15 Husheng Han , Xinyao Zheng , Yuanbo Wen , Yifan Hao , Erhu Feng , Ling Liang , Jianan Mu , Xiaqing Li , Tianyun Ma , Pengwei Jin , Xinkai Song , Zidong Du , Qi Guo , Xing Hu

Disaggregated systems have a novel architecture motivated by the requirements of resource intensive applications such as social networking, search, and in-memory databases. The total amount of resources such as memory and CPU cores is very…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-03 Ewnetu Bayuh Lakew , Petter Svärd , Erik Elmroth , Johan Tordsson

Computational workflows are a common class of application on supercomputers, yet the loosely coupled and heterogeneous nature of workflows often fails to take full advantage of their capabilities. We created Colmena to leverage the massive…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-27 Logan Ward , J. Gregory Pauloski , Valerie Hayot-Sasson , Yadu Babuji , Alexander Brace , Ryan Chard , Kyle Chard , Rajeev Thakur , Ian Foster

The Workflows Community Summit gathered 111 participants from 18 countries to discuss emerging trends and challenges in scientific workflows, focusing on six key areas: time-sensitive workflows, AI-HPC convergence, multi-facility workflows,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-22 Rafael Ferreira da Silva , Deborah Bard , Kyle Chard , Shaun de Witt , Ian T. Foster , Tom Gibbs , Carole Goble , William Godoy , Johan Gustafsson , Utz-Uwe Haus , Stephen Hudson , Shantenu Jha , Laila Los , Drew Paine , Frédéric Suter , Logan Ward , Sean Wilkinson , Marcos Amaris , Yadu Babuji , Jonathan Bader , Riccardo Balin , Daniel Balouek , Sarah Beecroft , Khalid Belhajjame , Rajat Bhattarai , Wes Brewer , Paul Brunk , Silvina Caino-Lores , Henri Casanova , Daniela Cassol , Jared Coleman , Taina Coleman , Iacopo Colonnelli , Anderson Andrei Da Silva , Daniel de Oliveira , Pascal Elahi , Nour Elfaramawy , Wael Elwasif , Brian Etz , Thomas Fahringer , Wesley Ferreira , Rosa Filgueira , Jacob Fosso Tande , Luiz Gadelha , Andy Gallo , Daniel Garijo , Yiannis Georgiou , Philipp Gritsch , Patricia Grubel , Amal Gueroudji , Quentin Guilloteau , Carlo Hamalainen , Rolando Hong Enriquez , Lauren Huet , Kevin Hunter Kesling , Paula Iborra , Shiva Jahangiri , Jan Janssen , Joe Jordan , Sehrish Kanwal , Liliane Kunstmann , Fabian Lehmann , Ulf Leser , Chen Li , Peini Liu , Jakob Luettgau , Richard Lupat , Jose M. Fernandez , Ketan Maheshwari , Tanu Malik , Jack Marquez , Motohiko Matsuda , Doriana Medic , Somayeh Mohammadi , Alberto Mulone , John-Luke Navarro , Kin Wai Ng , Klaus Noelp , Bruno P. Kinoshita , Ryan Prout , Michael R. Crusoe , Sashko Ristov , Stefan Robila , Daniel Rosendo , Billy Rowell , Jedrzej Rybicki , Hector Sanchez , Nishant Saurabh , Sumit Kumar Saurav , Tom Scogland , Dinindu Senanayake , Woong Shin , Raul Sirvent , Tyler Skluzacek , Barry Sly-Delgado , Stian Soiland-Reyes , Abel Souza , Renan Souza , Domenico Talia , Nathan Tallent , Lauritz Thamsen , Mikhail Titov , Benjamin Tovar , Karan Vahi , Eric Vardar-Irrgang , Edite Vartina , Yuandou Wang , Merridee Wouters , Qi Yu , Ziad Al Bkhetan , Mahnoor Zulfiqar

Job schedulers are a key component of scalable computing infrastructures. They orchestrate all of the work executed on the computing infrastructure and directly impact the effectiveness of the system. Recently, job workloads have…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-03-06 Albert Reuther , Chansup Byun , William Arcand , David Bestor , Bill Bergeron , Matthew Hubbell , Michael Jones , Peter Michaleas , Andrew Prout , Antonio Rosa , Jeremy Kepner

In recent years, as the demand for low energy and high performance computing has steadily increased, heterogeneous computing has emerged as an important and promising solution. Because most workloads can typically run most efficiently on…

Performance · Computer Science 2017-12-11 Zhuo Chen , Diana Marculescu

Deep research agents, which synthesize information across diverse sources, are significantly constrained by the sequential nature of reasoning. This bottleneck results in high latency, poor runtime adaptability, and inefficient resource…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-31 Lunyiu Nie , Nedim Lipka , Ryan A. Rossi , Swarat Chaudhuri