English
Related papers

Related papers: A Sensitivity Analysis of Flexibility from GPU-Hea…

200 papers

The use of High Performance Computing (HPC) in commercial and consumer IT applications is becoming popular. They need the ability to gain rapid and scalable access to high-end computing capabilities. Cloud computing promises to deliver such…

Distributed, Parallel, and Cluster Computing · Computer Science 2009-09-08 Saurabh Kumar Garg , Chee Shin Yeo , Arun Anandasivam , Rajkumar Buyya

The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-16 Jiaang Duan , Shenglin Xu , Shiyou Qian , Dingyu Yang , Kangjin Wang , Chenzhi Liao , Yinghao Yu , Qin Hua , Hanwen Hu , Qi Wang , Wenchao Wu , Dongqing Bao , Tianyu Lu , Jian Cao , Guangtao Xue , Guodong Yang , Liping Zhang , Gang Chen

Data centers (DCs) are emerging as large, geographically distributed, controllable loads whose participation in electricity markets can significantly affect grid operation, especially when cloud platforms shift workloads across sites to…

Systems and Control · Electrical Eng. & Systems 2026-04-09 Shijie Pan , Zaint A. Alexakis , Charalambos Konstantinou

Optimizing resource utilization in high-performance computing (HPC) clusters is essential for maximizing both system efficiency and user satisfaction. However, traditional rigid job scheduling often results in underutilized resources and…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-20 Patrick Zojer , Jonas Posner , Taylan Özden

We formulate optimization problems to study how data centers might modulate their power demands for cost-effective operation taking into account three key complex features exhibited by real-world electricity pricing schemes: (i)…

Systems and Control · Computer Science 2013-09-05 Cheng Wang , Bhuvan Urgaonkar , Qian Wang , George Kesidis , Anand Sivasubramaniam

Modern GPU datacenters are critical for delivering Deep Learning (DL) models and services in both the research community and industry. When operating a datacenter, optimization of resource scheduling and management can bring significant…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-09-07 Qinghao Hu , Peng Sun , Shengen Yan , Yonggang Wen , Tianwei Zhang

Recent research shows large-scale AI-centric data centers could experience rapid fluctuations in power demand due to varying computation loads, such as sudden spikes from inference or interruption of training large language models (LLMs).…

Signal Processing · Electrical Eng. & Systems 2025-03-12 Mariam Mughees , Yuzhuo Li , Yize Chen , Yunwei Ryan Li

Large scale-free graphs are famously difficult to process efficiently: the skewed vertex degree distribution makes it difficult to obtain balanced partitioning. Our research instead aims to turn this into an advantage by partitioning the…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-10-05 Scott Sallinen , Abdullah Gharaibeh , Matei Ripeanu

Data centers consume a large amount of energy and incur substantial electricity cost. In this paper, we study the familiar problem of reducing data center energy cost with two new perspectives. First, we find, through an empirical study of…

Networking and Internet Architecture · Computer Science 2013-12-18 Hong Xu , Baochun Li

Major innovations in computing have been driven by scaling up computing infrastructure, while aggressively optimizing operating costs. The result is a network of worldwide datacenters that consume a large amount of energy, mostly in an…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-30 Walid A. Hanafy , Roozbeh Bostandoost , Noman Bashir , David Irwin , Mohammad Hajiesmaili , Prashant Shenoy

Training Deep Neural Networks (DNNs) is a widely popular workload in both enterprises and cloud data centers. Existing schedulers for DNN training consider GPU as the dominant resource, and allocate other resources such as CPU and memory…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-08-25 Jayashree Mohan , Amar Phanishayee , Janardhan Kulkarni , Vijay Chidambaram

Renewable Energy Sources play a key role in smart energy systems. To achieve 100% renewable energy, utilizing the flexibility potential on the demand side becomes the cost-efficient option to balance the grid. However, it is not trivial to…

Systems and Control · Electrical Eng. & Systems 2024-08-01 Seyed Shahabaldin Tohidi , Henrik Madsen , Davide Calì , Tobias K. S. Ritschel

Enabling continued data-center growth under increasing grid stress motivates closer coordination between flexible computing demand and co-located battery energy storage systems (BESS) to improve site operations and provide grid services.…

Systems and Control · Electrical Eng. & Systems 2026-05-18 Shaohui Liu , Sungho Shin , Deepjyoti Deka

While the rapid expansion of data centers poses challenges for power grids, it also offers new opportunities as potentially flexible loads. Existing power system research often abstracts data centers as aggregate resources, while computer…

Systems and Control · Electrical Eng. & Systems 2026-02-06 Zhirui Liang , Jae-Won Chung , Mosharaf Chowdhury , Jiasi Chen , Vladimir Dvorkin

Energy costs are a major factor in the total cost of ownership (TCO) for high-performance computing (HPC) systems. The rise of intermittent green energy sources and reduced reliance on fossil fuels have introduced volatility into…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-02 Peter Arzt , Felix Wolf

Due to their highly parallel multi-cores architecture, GPUs are being increasingly used in a wide range of computationally intensive applications. Compared to CPUs, GPUs can achieve higher performances at accelerating the programs'…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-10-05 Frédéric Magoulès , Abal-Kassim Cheik Ahamed , Alban Desmaison , Jean-Christophe Léchenet , François Mayer , Haifa Ben Salem , Thomas Zhu

In this paper, we design an analytically and experimentally better online energy and job scheduling algorithm with the objective of maximizing net profit for a service provider in green data centers. We first study the previously known…

Performance · Computer Science 2014-04-22 Huangxin Wang , Jean X. Zhang , Fei Li

The ever-increasing computation and energy demand for LLM and AI agents call for holistic and efficient optimization of LLM serving systems. In practice, heterogeneous GPU clusters can be deployed in a geographically distributed manner,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-06 Xuan He , Zequan Fang , Jinzhao Lian , Danny H. K. Tsang , Baosen Zhang , Yize Chen

Further electrification of the economy is expected to sharpen ramp rates and increase peak loads. Flexibility from the demand side, which new technologies might facilitate, can help these operational challenges. Electric utilities have…

Systems and Control · Electrical Eng. & Systems 2025-07-10 Lane D. Smith , Daniel S. Kirschen

With the increasing popularity of Internet-based services and applications, power efficiency is becoming a major concern for data center operators, as high electricity consumption not only increases greenhouse gas emissions, but also…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-03-19 Dmytro Dyachuk , Michele Mazzucco