English
Related papers

Related papers: Cultivating Multidisciplinary AI Workforce Develop…

200 papers

Ensuring the highest training throughput to maximize resource efficiency, while maintaining fairness among users, is critical for deep learning (DL) training in heterogeneous GPU clusters. However, current DL schedulers provide only limited…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-03-28 Zizhao Mo , Huanle Xu , Wing Cheong Lau

Implementing existing federated learning in massive Internet of Things (IoT) networks faces critical challenges such as imbalanced and statistically heterogeneous data and device diversity. To this end, we propose a semi-federated learning…

Machine Learning · Computer Science 2023-03-10 Wanli Ni , Jingheng Zheng , Hui Tian

Generative artificial intelligence (GenAI) offers various services to users through content creation, which is believed to be one of the most important components in future networks. However, training and deploying big artificial…

Networking and Internet Architecture · Computer Science 2024-01-04 Yuqing Tian , Zhaoyang Zhang , Yuzhi Yang , Zirui Chen , Zhaohui Yang , Richeng Jin , Tony Q. S. Quek , Kai-Kit Wong

GPUs are vastly underutilized, even when running resource-intensive AI applications, as GPU kernels within each job have diverse resource profiles that may saturate some parts of a device while often leaving other parts idle. Colocating…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-17 Paul Elvinger , Foteini Strati , Natalie Enright Jerger , Ana Klimovic

The foundation-model ecosystem remains highly centralized because training requires immense compute resources and is therefore largely limited to large cloud operators. Edge-assisted foundation model training that harnesses spare compute on…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-14 Leyang Xue , Meghana Madhyastha , Myungjin Lee , Amos Storkey , Randal Burns , Mahesh K. Marina

Graph neural networks (GNNs) are one of the rapidly growing fields within deep learning. While many distributed GNN training frameworks have been proposed to increase the training throughput, they face three limitations when applied to…

Machine Learning · Computer Science 2024-08-14 Jaeyong Song , Hongsun Jang , Jaewon Jung , Youngsok Kim , Jinho Lee

The radical advances in mobile computing, the IoT technological evolution along with cyberphysical components (e.g., sensors, actuators, control centers) have led to the development of smart city applications that generate raw or…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-25 Dimitrios Tomaras , Michail Tsenos , Vana Kalogeraki , Dimitrios Gunopulos

GPUs are essential to accelerating the latency-sensitive deep neural network (DNN) inference workloads in cloud datacenters. To fully utilize GPU resources, spatial sharing of GPUs among co-located DNN inference workloads becomes…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-11-04 Fei Xu , Jianian Xu , Jiabin Chen , Li Chen , Ruitao Shang , Zhi Zhou , Fangming Liu

The advances in data, computing and networking over the last two decades led to a shift in many application domains that includes machine learning on big data as a part of the scientific process, requiring new capabilities for integrated…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-03-19 Ilkay Altintas , Kyle Marcus , Isaac Nealey , Scott L. Sellars , John Graham , Dima Mishin , Joel Polizzi , Daniel Crawl , Thomas DeFanti , Larry Smarr

There is an ever-increasing need for computational power to train complex artificial intelligence (AI) & machine learning (ML) models to tackle large scientific problems. High performance computing (HPC) resources are required to…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-05-22 David Brayford , Sofia Vallercorsa

Foundation models are at the forefront of AI research, appealing for their ability to learn from vast datasets and cater to diverse tasks. Yet, their significant computational demands raise issues of environmental impact and the risk of…

Machine Learning · Computer Science 2025-07-03 Leyang Xue , Meghana Madhyastha , Randal Burns , Myungjin Lee , Mahesh K. Marina

Training large-scale deep learning models has become a key challenge for the scientific community and industry. While the massive use of GPUs can significantly speed up training times, this approach has a negative impact on efficiency. In…

Machine Learning · Computer Science 2025-09-04 David Cortes , Carlos Juiz , Belen Bermejo

Training transformer models requires substantial GPU compute and memory resources. In homogeneous clusters, distributed strategies allocate resources evenly, but this approach is inefficient for heterogeneous clusters, where GPUs differ in…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-11-15 Runsheng Benson Guo , Utkarsh Anand , Arthur Chen , Khuzaima Daudjee

As we approach the Exascale era, it is important to verify that the existing frameworks and tools will still work at that scale. Moreover, public Cloud computing has been emerging as a viable solution for both prototyping and urgent…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-06-23 Igor Sfiligoi , Frank Wuerthwein , Benedikt Riedel , David Schultz

Significant investments to upgrade and construct large-scale scientific facilities demand commensurate investments in R&D to design algorithms and computing approaches to enable scientific and engineering breakthroughs in the big data era.…

Training GUI agents with traditional centralized methods faces significant cost and scalability challenges. Federated learning (FL) offers a promising solution, yet its potential is hindered by the lack of benchmarks that capture…

Multiagent Systems · Computer Science 2026-04-17 Wenhao Wang , Haoting Shi , Mengying Yuan , Yiquan Lin , Panrong Tong , Hanzhang Zhou , Guangyi Liu , Pengxiang Zhao , Yue Wang , Siheng Chen

Artificial intelligence (AI) is fueling exponential electricity demand growth, threatening grid reliability, raising prices for communities paying for new energy infrastructure, and stunting AI innovation as data centers wait for…

GPU-based heterogeneous architectures are now commonly used in HPC clusters. Due to their architectural simplicity specialized for data-level parallelism, GPUs can offer much higher computational throughput and memory bandwidth than CPUs in…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-05-15 Urvij Saroliya , Eishi Arima , Dai Liu , Martin Schulz

Federated Learning (FL) enables training Artificial Intelligence (AI) models over end devices without compromising their privacy. As computing tasks are increasingly performed by a combination of cloud, edge, and end devices, FL can benefit…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-30 Zhiyuan Wu , Sheng Sun , Yuwei Wang , Min Liu , Bo Gao , Quyang Pan , Tianliu He , Xuefeng Jiang

High-Performance Computing (HPC) centers and cloud providers support an increasingly diverse set of applications on heterogenous hardware. As Artificial Intelligence (AI) and Machine Learning (ML) workloads have become an increasingly…

‹ Prev 1 3 4 5 6 7 10 Next ›