English
Related papers

Related papers: The MIT Supercloud Workload Classification Challen…

200 papers

HPC systems used for research run a wide variety of software and workflows. This software is often written or modified by users to meet the needs of their research projects, and rarely is built with security in mind. In this paper we…

Emerging multi-model workloads with heavy models like recent large language models significantly increased the compute and memory demands on hardware. To address such increasing demands, designing a scalable hardware architecture became a…

Hardware Architecture · Computer Science 2024-09-17 Mohanad Odema , Luke Chen , Hyoukjun Kwon , Mohammad Abdullah Al Faruque

At present there are a number of barriers to creating an energy efficient workload scheduler for a Private Cloud based data center. Firstly, the relationship between different workloads and power consumption must be investigated. Secondly,…

Distributed, Parallel, and Cluster Computing · Computer Science 2011-05-16 James W. Smith , Ian Sommerville

AI workloads, particularly those driven by deep learning, are introducing novel usage patterns to high-performance computing (HPC) systems that are not comprehensively captured by standard HPC benchmarks. As one of the largest academic…

With the growing amount of data, data processing workloads and the management of their resource usage becomes increasingly important. Since managing a dedicated infrastructure is in many situations infeasible or uneconomical, users…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-01-19 Dominik Scheinert , Alireza Alamgiralem , Jonathan Bader , Jonathan Will , Thorsten Wittkopp , Lauritz Thamsen

Modern cloud-native systems increasingly rely on multi-cluster deployments to support scalability, resilience, and geographic distribution. However, existing resource management approaches remain largely reactive and cluster-centric,…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-01 Vinoth Punniyamoorthy , Akash Kumar Agarwal , Bikesh Kumar , Abhirup Mazumder , Kabilan Kannan , Sumit Saha

Understanding inter-VM interference is of paramount importance to provide a sound knowledge and understand where performance degradation comes from in the current public cloud. With this aim, this paper devises a workload taxonomy that…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-13 Lucia Pons , Josué Feliu , José Puche , Chaoyi Huang , Salvador Petit , Julio Pons , María E. Gómez , Julio Sahuquillo

Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture the real execution statistics…

Artificial Intelligence (AI) and Internet of Things (IoT) applications are rapidly growing in today's world where they are continuously connected to the internet and process, store and exchange information among the devices and the…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-05-01 Saravanan Ramanathan , Nitin Shivaraman , Seima Suryasekaran , Arvind Easwaran , Etienne Borde , Sebastian Steinhorst

The Workflows Community Summit gathered 111 participants from 18 countries to discuss emerging trends and challenges in scientific workflows, focusing on six key areas: time-sensitive workflows, AI-HPC convergence, multi-facility workflows,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-22 Rafael Ferreira da Silva , Deborah Bard , Kyle Chard , Shaun de Witt , Ian T. Foster , Tom Gibbs , Carole Goble , William Godoy , Johan Gustafsson , Utz-Uwe Haus , Stephen Hudson , Shantenu Jha , Laila Los , Drew Paine , Frédéric Suter , Logan Ward , Sean Wilkinson , Marcos Amaris , Yadu Babuji , Jonathan Bader , Riccardo Balin , Daniel Balouek , Sarah Beecroft , Khalid Belhajjame , Rajat Bhattarai , Wes Brewer , Paul Brunk , Silvina Caino-Lores , Henri Casanova , Daniela Cassol , Jared Coleman , Taina Coleman , Iacopo Colonnelli , Anderson Andrei Da Silva , Daniel de Oliveira , Pascal Elahi , Nour Elfaramawy , Wael Elwasif , Brian Etz , Thomas Fahringer , Wesley Ferreira , Rosa Filgueira , Jacob Fosso Tande , Luiz Gadelha , Andy Gallo , Daniel Garijo , Yiannis Georgiou , Philipp Gritsch , Patricia Grubel , Amal Gueroudji , Quentin Guilloteau , Carlo Hamalainen , Rolando Hong Enriquez , Lauren Huet , Kevin Hunter Kesling , Paula Iborra , Shiva Jahangiri , Jan Janssen , Joe Jordan , Sehrish Kanwal , Liliane Kunstmann , Fabian Lehmann , Ulf Leser , Chen Li , Peini Liu , Jakob Luettgau , Richard Lupat , Jose M. Fernandez , Ketan Maheshwari , Tanu Malik , Jack Marquez , Motohiko Matsuda , Doriana Medic , Somayeh Mohammadi , Alberto Mulone , John-Luke Navarro , Kin Wai Ng , Klaus Noelp , Bruno P. Kinoshita , Ryan Prout , Michael R. Crusoe , Sashko Ristov , Stefan Robila , Daniel Rosendo , Billy Rowell , Jedrzej Rybicki , Hector Sanchez , Nishant Saurabh , Sumit Kumar Saurav , Tom Scogland , Dinindu Senanayake , Woong Shin , Raul Sirvent , Tyler Skluzacek , Barry Sly-Delgado , Stian Soiland-Reyes , Abel Souza , Renan Souza , Domenico Talia , Nathan Tallent , Lauritz Thamsen , Mikhail Titov , Benjamin Tovar , Karan Vahi , Eric Vardar-Irrgang , Edite Vartina , Yuandou Wang , Merridee Wouters , Qi Yu , Ziad Al Bkhetan , Mahnoor Zulfiqar

Cloud computing has revolutionized the provisioning of computing resources, offering scalable, flexible, and on-demand services to meet the diverse requirements of modern applications. At the heart of efficient cloud operations are job…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-03 Yan Gu , Zhaoze Liu , Shuhong Dai , Cong Liu , Ying Wang , Shen Wang , Georgios Theodoropoulos , Long Cheng

High-performance computing (HPC) systems increasingly support both scalable AI training and large-scale simulation workloads. Both typically rely heavily on collective communication operations. On modern supercomputers, however, network…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-14 Lorenzo Piarulli , Marco Faltelli , Dirk Pleiter , Karthee Sivalingam , Dancheng Zhang , Kexue Zhao , Matteo Turisini , Francesco Iannone , Aldo Artigiani , Daniele De Sensi

Many organizations routinely analyze large datasets using systems for distributed data-parallel processing and clusters of commodity resources. Yet, users need to configure adequate resources for their data processing jobs. This requires…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-06-02 Lauritz Thamsen , Dominik Scheinert , Jonathan Will , Jonathan Bader , Odej Kao

Since emerging edge applications such as Internet of Things (IoT) analytics and augmented reality have tight latency constraints, hardware AI accelerators have been recently proposed to speed up deep neural network (DNN) inference run by…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-01-20 Qianlin Liang , Walid A. Hanafy , Ahmed Ali-Eldin , Prashant Shenoy

Several companies and research institutes are moving their CPU-intensive applications to hybrid High Performance Computing (HPC) cloud environments. Such a shift depends on the creation of software systems that help users decide where a job…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-08-30 Renato L. F. Cunha , Eduardo R. Rodrigues , Leonardo P. Tizzei , Marco A. S. Netto

The scale of scientific High Performance Computing (HPC) and High Throughput Computing (HTC) has increased significantly in recent years, and is becoming sensitive to total energy use and cost. Energy-efficiency has thus become an important…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-10-14 David Abdurachmanov , Peter Elmer , Giulio Eulisse , Robert Knight , Tapio Niemi , Jukka K. Nurminen , Filip Nyback , Goncalo Pestana , Zhonghong Ou , Kashif Khan

Long-running service workloads (e.g. web search engine) and short-term data analysis workloads (e.g. Hadoop MapReduce jobs) co-locate in today's data centers. Developing realistic benchmarks to reflect such practical scenario of mixed…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-12-07 Rui Han , Shulin Zhan , Chenrong Shao , Junwei Wang , Lizy K. John , Jiangtao Xu , Gang Lu , Lei Wang

Output-intensive scientific applications are highly sensitive to low storage throughput. While existing scientific application stacks are optimized for traditional High-Performance Computing (HPC) environments with high remote storage and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-22 Steven W. D. Chien , Kento Sato , Artur Podobas , Niclas Jansson , Stefano Markidis , Michio Honda

The integration of generative AI models, particularly large language models (LLMs), into real-time multi-model AI applications such as video conferencing and gaming is giving rise to a new class of workloads: real-time generative AI…

Machine Learning · Computer Science 2025-07-22 Rachid Karami , Rajeev Patwari , Hyoukjun Kwon , Ashish Sirasao

Modern AI clusters, which host diverse workloads like data pre-processing, training and inference, often store the large-volume data in cloud storage and employ caching frameworks to facilitate remote data access. To avoid code-intrusion…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-17 Tianze Wang , Yifei Liu , Chen Chen , Pengfei Zuo , Jiawei Zhang , Qizhen Weng , Yin Chen , Zhenhua Han , Jieru Zhao , Quan Chen , Minyi Guo