MIT Supercloud 数据集
分布式、并行与集群计算
2021-08-05 v1 人工智能
机器学习
摘要
人工智能(AI)和机器学习(ML)工作负载在传统高性能计算(HPC)中心和商业云系统中的计算工作负载占比日益增大。这导致了 HPC 集群和商业云部署方式的变化,以及新的重点放在优化资源使用、分配和部署新 AI 框架,以及诸如 Jupyter notebooks 等支持快速原型设计和部署的能力上。随着这些变化,需要更好地理解集群/数据中心运营,以开发改进的调度策略、识别资源利用中的低效、能耗/功耗、故障预测以及识别策略违规。本文介绍 MIT Supercloud 数据集,旨在促进用于大规模 HPC 和数据中心/云运营分析的新型 AI/ML 方法。我们提供了来自 MIT Supercloud 系统的详细监控日志,包括作业级别的 CPU 和 GPU 使用率、内存使用、文件系统日志以及物理监控数据。本文讨论了该数据集的细节、收集方法、数据可用性,并探讨了利用该数据正在开发的潜在挑战问题。数据集及未来的挑战公告将通过 https://dcc.mit.edu 提供。
引用
@article{arxiv.2108.02037,
title = {The MIT Supercloud Dataset},
author = {Siddharth Samsi and Matthew L Weiss and David Bestor and Baolin Li and Michael Jones and Albert Reuther and Daniel Edelman and William Arcand and Chansup Byun and John Holodnack and Matthew Hubbell and Jeremy Kepner and Anna Klein and Joseph McDonald and Adam Michaleas and Peter Michaleas and Lauren Milechin and Julia Mullen and Charles Yee and Benjamin Price and Andrew Prout and Antonio Rosa and Allan Vanterpool and Lindsey McEvoy and Anson Cheng and Devesh Tiwari and Vijay Gadepally},
journal= {arXiv preprint arXiv:2108.02037},
year = {2021}
}