在HPC集群中服务LLM:对比性研究——Qualcomm Cloud AI 100 Ultra与NVIDIA数据中心GPU
分布式、并行与集群计算
2025-10-30 v3 人工智能
摘要
本研究对Qualcomm Cloud AI 100 Ultra(QAic)加速器用于大型语言模型(LLM)推理进行基准测试分析,评估其能源效率(吞吐量/瓦特)、性能和硬件可扩展性,与NVIDIA A100 GPU(在4x和8x配置中)进行比较,置于National Research Platform(NRP)生态系统中。共服务12个开源LLM,参数规模从1.24亿到70亿不等,使用vLLM框架。我们的分析发现,QAic在特定模型上实现竞争性能源效率,同时实现更细粒度的硬件分配:部分70B模型仅需1块QAic即可运行,而需要8个A100 GPU,功耗降低20倍(148W vs 2983W)。对于较小的模型,单个QAic设备相较于我们的4 GPU A100配置功耗降低最高可达35倍(36W vs 1246W)。这些发现为Qualcomm Cloud AI 100 Ultra在能源受限和资源高效的HPC部署中潜力提供了见解。
关键词
引用
@article{arxiv.2507.00418,
title = {Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs},
author = {Mohammad Firas Sada and John J. Graham and Elham E Khoda and Mahidhar Tatineni and Dmitry Mishin and Rajesh K. Gupta and Rick Wagner and Larry Smarr and Thomas A. DeFanti and Frank Würthwein},
journal= {arXiv preprint arXiv:2507.00418},
year = {2025}
}
备注
8 pages, 3 tables