English

Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs

Distributed, Parallel, and Cluster Computing 2021-12-13 v1

Abstract

Diverse workloads such as interactive supercomputing, big data analysis, and large-scale AI algorithm development, requires a high-performance scheduler. This paper presents a novel node-based scheduling approach for large scale simulations of short running jobs on MIT SuperCloud systems, that allows the resources to be fully utilized for both long running batch jobs while simultaneously providing fast launch and release of large-scale short running jobs. The node-based scheduling approach has demonstrated up to 100 times faster scheduler performance that other state-of-the-art systems.

Keywords

Cite

@article{arxiv.2108.11359,
  title  = {Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs},
  author = {Chansup Byun and William Arcand and David Bestor and Bill Bergeron and Vijay Gadepally and Michael Houle and Matthew Hubbell and Michael Jones and Anna Klein and Peter Michaleas and Lauren Milechin and Julie Mullen and Andrew Prout and Albert Reuther and Antonio Rosa and Siddharth Samsi and Charles Yee and Jeremy Kepner},
  journal= {arXiv preprint arXiv:2108.11359},
  year   = {2021}
}

Comments

IEEE HPEC 2021