English

ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving

Distributed, Parallel, and Cluster Computing 2024-07-16 v1 Artificial Intelligence

Abstract

Since the increasing popularity of large language model (LLM) backend systems, it is common and necessary to deploy stable serverless serving of LLM on multi-GPU clusters with autoscaling. However, there exist challenges because the diversity and co-location of applications in multi-GPU clusters will lead to low service quality and GPU utilization. To address them, we build ENOVA, a deployment, monitoring and autoscaling service towards serverless LLM serving. ENOVA deconstructs the execution process of LLM service comprehensively, based on which ENOVA designs a configuration recommendation module for automatic deployment on any GPU clusters and a performance detection module for autoscaling. On top of them, ENOVA implements a deployment execution engine for multi-GPU cluster scheduling. The experiment results show that ENOVA significantly outperforms other state-of-the-art methods and is suitable for wide deployment in large online systems.

Keywords

Cite

@article{arxiv.2407.09486,
  title  = {ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving},
  author = {Tao Huang and Pengfei Chen and Kyoka Gong and Jocky Hawk and Zachary Bright and Wenxin Xie and Kecheng Huang and Zhi Ji},
  journal= {arXiv preprint arXiv:2407.09486},
  year   = {2024}
}
R2 v1 2026-06-28T17:39:02.508Z