English

Design and Operation of Shared Machine Learning Clusters on Campus

Distributed, Parallel, and Cluster Computing 2025-05-15 v2 Networking and Internet Architecture

Abstract

Amid the rapid advancements in large machine learning (ML) models, universities worldwide are investing substantial funds and efforts into GPU clusters. However, managing a shared GPU cluster poses a pyramid of challenges, from hardware configuration to resource allocation among users. This paper introduces SING, a full-stack solution designed to streamline the management of shared GPU clusters in academic institutions. Motivated by the pressing need for efficient resource sharing and the challenges posed by limited staffing, we present a comprehensive view of SING's architecture and design choices, which achieves operational efficiency (i.e., low maintenance cost and high resource utilization). We also share experience and insights from the real-world operations of SING, including analysis of its usage patterns and management of incidents and failures. This paper is part of our ongoing effort to improve the management of shared ML clusters. We open-source relevant resources to facilitate the development and operation of similar clusters for ML.

Keywords

Cite

@article{arxiv.2110.01556,
  title  = {Design and Operation of Shared Machine Learning Clusters on Campus},
  author = {Kaiqiang Xu and Decang Sun and Hao Wang and Zhenghang Ren and Xinchen Wan and Xudong Liao and Zilong Wang and Junxue Zhang and Kai Chen},
  journal= {arXiv preprint arXiv:2110.01556},
  year   = {2025}
}

Comments

Accepted to ACM ASPLOS 2025 (ACM International Conference on Architectural Support for Programming Languages and Operating Systems). Extended version

R2 v1 2026-06-24T06:36:44.650Z