Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
Abstract
The rapid progress in Deep Learning (DL) and Large Language Models (LLMs) has exponentially increased demands of computational power and bandwidth. This, combined with the high costs of faster computing chips and interconnects, has significantly inflated High Performance Computing (HPC) construction costs. To address these challenges, we introduce the Fire-Flyer AI-HPC architecture, a synergistic hardware-software co-design framework and its best practices. For DL training, we deployed the Fire-Flyer 2 with 10,000 PCIe A100 GPUs, achieved performance approximating the DGX-A100 while reducing costs by half and energy consumption by 40%. We specifically engineered HFReduce to accelerate allreduce communication and implemented numerous measures to keep our Computation-Storage Integrated Network congestion-free. Through our software stack, including HaiScale, 3FS, and HAI-Platform, we achieved substantial scalability by overlapping computation and communication. Our system-oriented experience from DL training provides valuable insights to drive future advancements in AI-HPC.
Cite
@article{arxiv.2408.14158,
title = {Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning},
author = {Wei An and Xiao Bi and Guanting Chen and Shanhuang Chen and Chengqi Deng and Honghui Ding and Kai Dong and Qiushi Du and Wenjun Gao and Kang Guan and Jianzhong Guo and Yongqiang Guo and Zhe Fu and Ying He and Panpan Huang and Jiashi Li and Wenfeng Liang and Xiaodong Liu and Xin Liu and Yiyuan Liu and Yuxuan Liu and Shanghao Lu and Xuan Lu and Xiaotao Nie and Tian Pei and Junjie Qiu and Hui Qu and Zehui Ren and Zhangli Sha and Xuecheng Su and Xiaowen Sun and Yixuan Tan and Minghui Tang and Shiyu Wang and Yaohui Wang and Yongji Wang and Ziwei Xie and Yiliang Xiong and Yanhong Xu and Shengfeng Ye and Shuiping Yu and Yukun Zha and Liyue Zhang and Haowei Zhang and Mingchuan Zhang and Wentao Zhang and Yichao Zhang and Chenggang Zhao and Yao Zhao and Shangyan Zhou and Shunfeng Zhou and Yuheng Zou},
journal= {arXiv preprint arXiv:2408.14158},
year = {2024}
}
Comments
This is the preprint version of the paper accepted for presentation at the 2024 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'24). \c{opyright} 2024 IEEE. Personal use of this material is permitted. For other uses, permission from IEEE must be obtained. Please refer to IEEE Xplore for the final published version