English

Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster

Hardware Architecture 2026-05-27 v2 Distributed, Parallel, and Cluster Computing Systems and Control Systems and Control

Abstract

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

Keywords

Cite

@article{arxiv.2605.24461,
  title  = {Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster},
  author = {Ehsan K. Ardestani and Leonardo Piga and Jovan Stojkovic and Pavan Balaji and Mustafa Ozdal and Mikel Jimenez Fernandez and Mihaela Dimovska and Luka Tadic and Hao Shen and Devika Vishwanath and Richa Mishra and Melaku Mihret and Valentin Andrei and Mauricio Cespedes and Julien Prigent and James Monahan and Tyler Graf and Bin Li and Charles Marquez and Shobhit Kanaujia and Kaushik Veeraraghavan and Chunqiang Tang},
  journal= {arXiv preprint arXiv:2605.24461},
  year   = {2026}
}
R2 v1 2026-07-22T07:29:51.740Z