English
Related papers

Related papers: Cultivating Multidisciplinary AI Workforce Develop…

200 papers

Grid computing has made substantial advances during the last decade. Grid middleware such as Globus has contributed greatly in making this possible. There are, however, significant barriers to the adoption of Grid computing in other fields,…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-11-17 Arshad Ali , Richard McClatchey , Ashiq Anjum , Irfan Habib , Kamran Soomro , Mohammed Asif , Ali Adil , Athar Mohsin

We present the design, implementation, and comprehensive evaluation of a specialized course on GPU architecture, GPU programming, and how these are used for developing AI agents. This course is offered to undergraduate and graduate students…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-18 Sriram Srinivasan , Hamdan Alabsi , Rand Obeidat , Nithisha Ponnala , Azene Zenebe

Accommodating long-running deep learning (DL) training and inference jobs is challenging on GPU clusters that use traditional batch schedulers, such as Slurm. Given fixed wall clock time limits, DL researchers usually need to run a sequence…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-27 Qiyang Ding , Pengfei Zheng , Shreyas Kudari , Shivaram Venkataraman , Zhao Zhang

In the ever evolving landscape of deep learning, unlocking the potential of cutting-edge models demands computational resources that surpass the capabilities of individual machines. Enter the NVIDIA DeepOps Slurm cluster, a meticulously…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-05-02 Arindam Majee

Over the past decades, progress in deployable autonomous flight systems has slowly stagnated. This is reflected in today's production air-crafts, where pilots only enable simple physics-based systems such as autopilot for takeoff, landing,…

Artificial Intelligence · Computer Science 2020-04-28 Andrew Wood , Ali Sydney , Peter Chin , Bishal Thapa , Ryan Ross

Grid is an infrastructure that involves the integrated and collaborative use of computers, networks, databases and scientific instruments owned and managed by multiple organizations. Grid applications often involve large amounts of data…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Parvin Asadzadeh , Rajkumar Buyya , Chun Ling Kei , Deepa Nayar , Srikumar Venugopal

Most existing training systems focus on a single region. In contrast, we envision that cross-region training offers more flexible GPU resource allocation and yields significant potential. However, the hierarchical cluster topology and…

Systems and Control · Electrical Eng. & Systems 2025-05-28 Jinquan Wang , Xiaojian Liao , Xuzhao Liu , Jiashun Suo , Zhisheng Huo , Chenhao Zhang , Xiangrong Xu , Runnan Shen , Xilong Xie , Limin Xiao

AI Infrastructure plays a key role in the speed and cost-competitiveness of developing and deploying advanced AI models. The current demand for powerful AI infrastructure for model training is driven by the emergence of generative AI and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-01-15 Talia Gershon , Seetharami Seelam , Brian Belgodere , Milton Bonilla , Lan Hoang , Danny Barnett , I-Hsin Chung , Apoorve Mohan , Ming-Hung Chen , Lixiang Luo , Robert Walkup , Constantinos Evangelinos , Shweta Salaria , Marc Dombrowa , Yoonho Park , Apo Kayi , Liran Schour , Alim Alim , Ali Sydney , Pavlos Maniotis , Laurent Schares , Bernard Metzler , Bengi Karacali-Akyamac , Sophia Wen , Tatsuhiro Chiba , Sunyanan Choochotkaew , Takeshi Yoshimura , Claudia Misale , Tonia Elengikal , Kevin O Connor , Zhuoran Liu , Richard Molina , Lars Schneidenbach , James Caden , Christopher Laibinis , Carlos Fonseca , Vasily Tarasov , Swaminathan Sundararaman , Frank Schmuck , Scott Guthridge , Jeremy Cohn , Marc Eshel , Paul Muench , Runyu Liu , William Pointer , Drew Wyskida , Bob Krull , Ray Rose , Brent Wolfe , William Cornejo , John Walter , Colm Malone , Clifford Perucci , Frank Franco , Nigel Hinds , Bob Calio , Pavel Druyan , Robert Kilduff , John Kienle , Connor McStay , Andrew Figueroa , Matthew Connolly , Edie Fost , Gina Roma , Jake Fonseca , Ido Levy , Michele Payne , Ryan Schenkel , Amir Malki , Lion Schneider , Aniruddha Narkhede , Shekeba Moshref , Alexandra Kisin , Olga Dodin , Bill Rippon , Henry Wrieth , John Ganci , Johnny Colino , Donna Habeger-Rose , Rakesh Pandey , Aditya Gidh , Aditya Gaur , Dennis Patterson , Samsuddin Salmani , Rambilas Varma , Rumana Rumana , Shubham Sharma , Aditya Gaur , Mayank Mishra , Rameswar Panda , Aditya Prasad , Matt Stallone , Gaoyuan Zhang , Yikang Shen , David Cox , Ruchir Puri , Dakshi Agrawal , Drew Thorstensen , Joel Belog , Brent Tang , Saurabh Kumar Gupta , Amitabha Biswas , Anup Maheshwari , Eran Gampel , Jason Van Patten , Matthew Runion , Sai Kaki , Yigal Bogin , Brian Reitz , Steve Pritko , Shahan Najam , Surya Nambala , Radhika Chirra , Rick Welp , Frank DiMitri , Felipe Telles , Amilcar Arvelo , King Chu , Ed Seminaro , Andrew Schram , Felix Eickhoff , William Hanson , Eric Mckeever , Michael Light , Dinakaran Joseph , Piyush Chaudhary , Piyush Shivam , Puneet Chaudhary , Wesley Jones , Robert Guthrie , Chris Bostic , Rezaul Islam , Steve Duersch , Wayne Sawdon , John Lewars , Matthew Klos , Michael Spriggs , Bill McMillan , George Gao , Ashish Kamra , Gaurav Singh , Marc Curry , Tushar Katarki , Joe Talerico , Zenghui Shi , Sai Sindhur Malleni , Erwan Gallen

Training large-scale models relies on a vast number of computing resources. For example, training the GPT-4 model (1.8 trillion parameters) requires 25000 A100 GPUs . It is a challenge to build a large-scale cluster with one type of…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-12 Si Xu , Zixiao Huang , Yan Zeng , Shengen Yan , Xuefei Ning , Quanlu Zhang , Haolin Ye , Sipei Gu , Chunsheng Shui , Zhezheng Lin , Hao Zhang , Sheng Wang , Guohao Dai , Yu Wang

In recent years, large language models have achieved great success due to their unprecedented size. However, training these models poses a challenge for most researchers as it requires a substantial number of GPUs. To reduce GPU memory…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-01 Haichen Huang , Jiarui Fang , Hongxin Liu , Shenggui Li , Yang You

Due to unfolded developments in both the IT sectors viz. Intelligent Transportation and Information Technology contemporary Smart Grid (SG) systems are leveraged with smart devices and entities. Such infrastructures when bestowed with the…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-02-07 Md. Muzakkir Hussain , Mohammad Saad Alam , M. M. Sufyan Beg

Large Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by…

Computation and Language · Computer Science 2025-04-29 Ranran Zhen , Juntao Li , Yixin Ji , Zhenlin Yang , Tong Liu , Qingrong Xia , Xinyu Duan , Zhefeng Wang , Baoxing Huai , Min Zhang

The use of GPUs has proliferated for machine learning workflows and is now considered mainstream for many deep learning models. Meanwhile, when training state-of-the-art personal recommendation models, which consume the highest number of…

Hardware Architecture · Computer Science 2020-11-12 Bilge Acun , Matthew Murphy , Xiaodong Wang , Jade Nie , Carole-Jean Wu , Kim Hazelwood

The surging development of Artificial Intelligence-Generated Content (AIGC) marks a transformative era of the content creation and production. Edge servers promise attractive benefits, e.g., reduced service delay and backhaul traffic load,…

Machine Learning · Computer Science 2024-09-10 Yuxin Liang , Peng Yang , Yuanyuan He , Feng Lyu

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies.…

Federated graph learning (FGL) enables multiple clients to collaboratively train powerful graph neural networks without sharing their private, decentralized graph data. Inherited from generic federated learning, FGL is critically challenged…

Machine Learning · Computer Science 2025-08-15 Xinrui Li , Qilin Fan , Tianfu Wang , Kaiwen Wei , Ke Yu , Xu Zhang

Autoregressive inference in large transformer-based language models (LLMs) presents significant challenges for runtime efficiency, particularly during the decode phase where load imbalance across GPU shards can cause throughput degradation…

Machine Learning · Computer Science 2025-09-24 Javed I. Khan an Henry Uwabor Moye

GPU clusters in multi-tenant settings often suffer from underutilization, making GPU-sharing technologies essential for efficient resource use. Among them, NVIDIA Multi-Instance GPU (MIG) has gained traction for providing hardware-level…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-14 Myeongsu Kim , Ikjun Yeom , Younghoon Kim

General-purpose Computing on Graphics Processing Units (GPGPU) has been introduced to many areas of scientific research such as bioinformatics, cryptography, computer vision, and deep learning. However, computing models in the High-energy…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-07-23 Max Isacson , Mattias Ellert , Richard Brenner

GPU singletasking is becoming increasingly inefficient and unsustainable as hardware capabilities grow and workloads diversify. We are now at an inflection point where GPUs must embrace multitasking, much like CPUs did decades ago, to meet…

Operating Systems · Computer Science 2025-08-13 Jiarong Xing , Yifan Qiao , Simon Mo , Xingqi Cui , Gur-Eyal Sela , Yang Zhou , Joseph Gonzalez , Ion Stoica