English
Related papers

Related papers: DeepServe: Serverless Large Language Model Serving…

200 papers

The high computational and memory requirements of generative large language models (LLMs) make it challenging to serve them cheaply. This paper aims to reduce the monetary cost for serving LLMs by leveraging preemptible GPU instances on…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-11-28 Xupeng Miao , Chunan Shi , Jiangfei Duan , Xiaoli Xi , Dahua Lin , Bin Cui , Zhihao Jia

Efficient execution of deep learning workloads on dataflow architectures is crucial for overcoming memory bottlenecks and maximizing performance. While streaming intermediate results between computation kernels can significantly improve…

Hardware Architecture · Computer Science 2025-09-24 Hanchen Ye , Deming Chen

The widespread adoption of cloud-based proprietary large language models (LLMs) has introduced significant challenges, including operational dependencies, privacy concerns, and the necessity of continuous internet connectivity. In this…

Machine Learning · Computer Science 2025-06-02 Chansung Park , Juyong Jiang , Fan Wang , Sayak Paul , Jing Tang

The past few years has witnessed specialized large language model (LLM) inference systems, such as vLLM, SGLang, Mooncake, and DeepFlow, alongside rapid LLM adoption via services like ChatGPT. Driving these system design efforts is the…

Databases · Computer Science 2025-06-30 James Pan , Guoliang Li

The Mixture of Experts (MoE) models are emerging as the latest paradigm for Large Language Models (LLMs). However, due to memory constraints, MoE models with billions or even trillions of parameters can only be deployed in multi-GPU or even…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-14 Bowen Zhou , Jinrui Jia , Wenhao He , Yong Zhang , Fang Dong

Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approaches often result in substantial inference slowdowns due to…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-19 Ruibo Fan , Xiangrui Yu , Xinglin Pan , Zeyu Li , Weile Luo , Qiang Wang , Wei Wang , Xiaowen Chu

Large Language Model (LLM) serving systems remain fundamentally fragile, where frequent hardware faults in hyperscale clusters trigger disproportionate service outages in the software stack. Current recovery mechanisms are prohibitively…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-02 Shangshu Qian , Kipling Liu , P. C. Sruthi , Lin Tan , Yongle Zhang

The rapid evolution of large language models (LLMs), driven by growing parameter scales, adoption of mixture-of-experts (MoE) architectures, and expanding context lengths, imposes unprecedented demands on AI infrastructure. Traditional AI…

Fine-tuning large language models (LLMs) on private, on-device data can empower tailored personalized AI agents. However, fine-tuning LLMs on resource-constrained edge devices faces significant challenges, including excessive computation…

Machine Learning · Computer Science 2025-03-26 Jian Ma , Xinchen Lyu , Jun Jiang , Qimei Cui , Haipeng Yao , Xiaofeng Tao

The traditional cloud-centric approach for Deep Learning (DL) requires training data to be collected and processed at a central server which is often challenging in privacy-sensitive domains like healthcare. Towards this, a new learning…

Cryptography and Security · Computer Science 2021-11-08 Andreas Grafberger , Mohak Chadha , Anshul Jindal , Jianfeng Gu , Michael Gerndt

Since the increasing popularity of large language model (LLM) backend systems, it is common and necessary to deploy stable serverless serving of LLM on multi-GPU clusters with autoscaling. However, there exist challenges because the…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-16 Tao Huang , Pengfei Chen , Kyoka Gong , Jocky Hawk , Zachary Bright , Wenxin Xie , Kecheng Huang , Zhi Ji

Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode phases. Existing approaches either (1) aggregate both phases…

Machine Learning · Computer Science 2025-11-10 Lei Gao , Chaoyi Jiang , Hossein Entezari Zarch , Daniel Wong , Murali Annavaram

Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks, but serving them efficiently at scale remains a critical challenge due to their substantial computational and latency demands. While most existing…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-04 Yifan Sun , Gholamreza Haffari , Minxian Xu , Rajkumar Buyya , Adel N. Toosi

Serverless computing is increasingly popular because of its lower cost and easier deployment. Several cloud service providers (CSPs) offer serverless computing on their public clouds, but it may bring the vendor lock-in risk. To avoid this…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-06-08 Junfeng Li , Sameer G. Kulkarni , K. K. Ramakrishnan , Dan Li

Large language model (LLM) serving infrastructures are undergoing a shift toward heterogeneity and disaggregation. Modern deployments increasingly integrate diverse accelerators and near-memory processing technologies, introducing…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-24 Jaehong Cho , Hyunmin Choi , Guseul Heo , Jongse Park

Large Language Models (LLMs) play a critical role in emerging agentic applications, where the timely completion of each entire inference is critical. Meanwhile, agentic LLM inferences are increasingly served on heterogeneous GPUs in…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-19 Boxiao Du , Boning Huangfu , Yizhou Luo , Chen Chen , Zijun Li , Minchen Yu , Xiaoyi Fan , Minyi Guo

Serverless computing has gained popularity in edge computing due to its flexible features, including the pay-per-use pricing model, auto-scaling capabilities, and multi-tenancy support. Complex Serverless-based applications typically rely…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-09-01 Ke Luo , Tao Ouyang , Zhi Zhou , Xu Chen

Dynamic offloading of Machine Learning (ML) model partitions across different resource orchestration services, such as Function-as-a-Service (FaaS) and Infrastructure-as-a-Service (IaaS), can balance processing and transmission delays while…

Machine Learning · Computer Science 2025-11-03 Zongshun Zhang , Ibrahim Matta

Large language models (LLMs) have revolutionized applications such as code completion, chatbots, and online classification. To elevate user experiences, service level objectives (SLOs) serve as crucial benchmarks for assessing inference…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-13 Jinqi Huang , Yi Xiong , Xuebing Yu , Wenjie Huang , Entong Li , Li Zeng , Xin Chen

Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentralized serving. This report presents xDeepServe, the…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-03 Ao Xiao , Bangzheng He , Baoquan Zhang , Baoxing Huai , Bingji Wang , Bo Wang , Bo Xu , Boyi Hou , Chan Yang , Changhong Liu , Cheng Cui , Chenyu Zhu , Cong Feng , Daohui Wang , Dayun Lin , Duo Zhao , Fengshao Zou , Fu Wang , Gangqiang Zhang , Gengyuan Dan , Guanjie Chen , Guodong Guan , Guodong Yang , Haifeng Li , Haipei Zhu , Haley Li , Hao Feng , Hao Huang , Hao Xu , Hengrui Ma , Hengtao Fan , Hui Liu , Jia Li , Jiang Liu , Jiang Xu , Jie Meng , Jinhan Xin , Junhao Hu , Juwei Chen , Lan Yu , Lanxin Miao , Liang Liu , Linan Jing , Lu Zhou , Meina Han , Mingkun Deng , Mingyu Deng , Naitian Deng , Nizhong Lin , Peihan Zhao , Peng Pan , Pengfei Shen , Ping Li , Qi Zhang , Qian Wang , Qin ZhC Qingrong Xia , Qingyi Zhang , Qunchao Fu , Ren Guo , Ruimin Gao , Shaochun Li , Sheng Long , Shentian Li , Shining Wan , Shuai Shen , Shuangfu Zeng , Shuming Jing , Siqi Yang , Song Zhang , Tao Xu , Tianlin Du , Ting Chen , Wanxu Wu , Wei Jiang , Weinan Tong , Weiwei Chen , Wen Peng , Wenli Zhou , Wenquan Yang , Wenxin Liang , Xiang Liu , Xiaoli Zhou , Xin Jin , Xinyu Duan , Xu Li , Xu Zhang , Xusheng Chen , Yalong Shan , Yang Gan , Yao Lu , Yi Deng , Yi Zheng , Ying Xiong , Yingfei Zheng , Yiyun Zheng , Yizhou Shan , Yong Gao , Yong Zhang , Yongqiang Yang , Yuanjin Gong , Yue Yu , Yuetao Chen , Yukun Zhu , Yulong He , Yusu Zhao , Yuyan Wu , Zenan Zhang , Zhaojin Zhuo , Zhaoyang Ji , Zhefeng Wang , Zheng Wang , Zhenan Fan , Zhenhua Yang , Zhenli Sheng , Zhibin Yu , Zhigang Ji , Zhihao Ren , Zhipeng Bian , Zhixia Liu , Zhiyu Dong , Zhonghua Li , Zhou Yu , Zhuoming Shen , Zhuwei Peng , Zi Ye , Zihao Xiang , Zimin Fu , Zixuan Zhang