English

Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Distributed, Parallel, and Cluster Computing 2024-07-30 v1

Abstract

Large Language Models (LLMs) like GPT and LLaMA are revolutionizing the AI industry with their sophisticated capabilities. Training these models requires vast GPU clusters and significant computing time, posing major challenges in terms of scalability, efficiency, and reliability. This survey explores recent advancements in training systems for LLMs, including innovations in training infrastructure with AI accelerators, networking, storage, and scheduling. Additionally, the survey covers parallelism strategies, as well as optimizations for computation, communication, and memory in distributed LLM training. It also includes approaches of maintaining system reliability over extended training periods. By examining current innovations and future directions, this survey aims to provide valuable insights towards improving LLM training systems and tackling ongoing challenges. Furthermore, traditional digital circuit-based computing systems face significant constraints in meeting the computational demands of LLMs, highlighting the need for innovative solutions such as optical computing and optical networks.

Keywords

Cite

@article{arxiv.2407.20018,
  title  = {Efficient Training of Large Language Models on Distributed Infrastructures: A Survey},
  author = {Jiangfei Duan and Shuo Zhang and Zerui Wang and Lijuan Jiang and Wenwen Qu and Qinghao Hu and Guoteng Wang and Qizhen Weng and Hang Yan and Xingcheng Zhang and Xipeng Qiu and Dahua Lin and Yonggang Wen and Xin Jin and Tianwei Zhang and Peng Sun},
  journal= {arXiv preprint arXiv:2407.20018},
  year   = {2024}
}
R2 v1 2026-06-28T17:56:54.834Z