面向短时作业大规模模拟的基于节点的作业调度
分布式、并行与集群计算
2021-12-13 v1
摘要
交互式超级计算、大数据分析以及大规模 AI 算法开发等多种工作负载都需要高性能调度器。本文提出了一种新颖的基于节点的调度方法,用于 MIT SuperCloud 系统上短时作业的大规模模拟,该方法允许资源被长时间运行的批处理作业充分利用,同时提供大规模短时作业的快速启动和释放。该基于节点的调度方法展现了比其他最先进系统快达 100 倍的调度性能。
引用
@article{arxiv.2108.11359,
title = {Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs},
author = {Chansup Byun and William Arcand and David Bestor and Bill Bergeron and Vijay Gadepally and Michael Houle and Matthew Hubbell and Michael Jones and Anna Klein and Peter Michaleas and Lauren Milechin and Julie Mullen and Andrew Prout and Albert Reuther and Antonio Rosa and Siddharth Samsi and Charles Yee and Jeremy Kepner},
journal= {arXiv preprint arXiv:2108.11359},
year = {2021}
}
备注
IEEE HPEC 2021