English
Related papers

Related papers: ReviveMoE: Fast Recovery for Hardware Failures in …

200 papers

Large Language Models (LLMs) can memorize sensitive information, raising concerns about potential misuse. LLM Unlearning, a post-hoc approach to remove this information from trained LLMs, offers a promising solution to mitigate these risks.…

Computation and Language · Computer Science 2024-09-19 Tianle Gu , Kexin Huang , Ruilin Luo , Yuanqi Yao , Yujiu Yang , Yan Teng , Yingchun Wang

Recent advancements in Large Language Models (LLMs) have significantly improved reasoning capabilities, with in-context learning (ICL) emerging as a key technique for adaptation without retraining. While previous works have focused on…

Machine Learning · Computer Science 2025-12-17 Jongyeop Hyun , Bumsoo Kim

Large language models (LLMs) have shown great potential in natural language processing and content generation. However, current LLMs heavily rely on cloud computing, leading to prolonged latency, high bandwidth cost, and privacy concerns.…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-05-24 Mingjin Zhang , Jiannong Cao , Xiaoming Shen , Zeyang Cui

Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different requests in the same batch to improve resource utilization…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-21 Amna Masood , Pratishtha Gaur , Nuwan Jayasena

To optimize the reasoning and problem-solving capabilities of Large Language Models (LLMs), we propose a novel cloud-edge collaborative architecture that enables a structured multi-agent prompting framework. This framework comprises three…

Computation and Language · Computer Science 2025-12-29 Shadikur Rahman , Aroosa Hameed , Gautam Srivastava , Syed Muhammad Danish

Large Language Model (LLM) inference services demand exceptionally high availability and low latency, yet multi-GPU Tensor Parallelism (TP) makes them vulnerable to single-GPU failures. We present AnchorTP, a state-preserving elastic TP…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-18 Wendong Xu , Chujie Chen , He Xiao , Kuan Li , Jing Xiong , Chen Zhang , Wenyong Zhou , Chaofan Tao , Yang Bai , Bei Yu , Ngai Wong

Mixture-of-Experts (MoE) based large language models (LLMs) offer strong performance but suffer from high memory and computation costs. Weight binarization provides extreme efficiency, yet existing binary methods designed for dense LLMs…

Machine Learning · Computer Science 2026-04-22 Zhixiong Zhao , Zukang Xu , Zhixuan Chen , Dawei Yang

In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also crucial for ML service providers as it helps lower the…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-01 Yunseong Kim , Yujeong Choi , Minsoo Rhu

Recent advances in large language models (LLMs) have enabled autonomous agents with complex reasoning and task-fulfillment capabilities using a wide range of tools. However, effectively identifying the most relevant tools for a given task…

Expert parallelism is vital for effectively training Mixture-of-Experts (MoE) models, enabling different devices to host distinct experts, with each device processing different input data. However, during expert parallel training, dynamic…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-13 Xinyi Liu , Yujie Wang , Fangcheng Fu , Xuefeng Xiao , Huixia Li , Jiashi Li , Bin Cui

Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges. 1) Ignoring the diversity (various prefixes and tidal requests), treating all the…

This paper presents MoE-Infinity, an efficient MoE inference system designed for personal machines with limited GPU memory capacity. The key idea for MoE-Infinity is that on personal machines, which are often single-user environments,…

Machine Learning · Computer Science 2025-03-14 Leyang Xue , Yao Fu , Zhan Lu , Luo Mai , Mahesh Marina

Mixture-of-Expert (MoE) presents a strong potential in enlarging the size of language model to trillions of parameters. However, training trillion-scale MoE requires algorithm and system co-design for a well-tuned high performance…

Machine Learning · Computer Science 2021-03-25 Jiaao He , Jiezhong Qiu , Aohan Zeng , Zhilin Yang , Jidong Zhai , Jie Tang

Deploying large language models (LLMs) on mobile devices is an emerging trend to enable data privacy and offline accessibility of LLM applications. Modern mobile neural processing units (NPUs) make such deployment increasingly feasible.…

Operating Systems · Computer Science 2026-04-13 Yongsheng Yan , Jiacheng Shen , Xuchuan Luo , Yangfan Zhou

Transformer based Large Language Models (LLMs) have been widely used in many fields, and the efficiency of LLM inference becomes hot topic in real applications. However, LLMs are usually complicatedly designed in model structure with…

Hardware Architecture · Computer Science 2024-06-25 Hui Wu , Yi Gan , Feng Yuan , Jing Ma , Wei Zhu , Yutao Xu , Hong Zhu , Yuhua Zhu , Xiaoli Liu , Jinghui Gu , Peng Zhao

Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-25 Li Wu , Walid A. Hanafy , Tarek Abdelzaher , David Irwin , Jesse Milzman , Prashant Shenoy

Deep Recommender Models (DLRMs) inference is a fundamental AI workload accounting for more than 79% of the total AI workload in Meta's data centers. DLRMs' performance bottleneck is found in the embedding layers, which perform many random…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-07-03 Giuseppe Ruggeri , Renzo Andri , Daniele Jahier Pagliari , Lukas Cavigelli

Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentralized serving. This report presents xDeepServe, the…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-03 Ao Xiao , Bangzheng He , Baoquan Zhang , Baoxing Huai , Bingji Wang , Bo Wang , Bo Xu , Boyi Hou , Chan Yang , Changhong Liu , Cheng Cui , Chenyu Zhu , Cong Feng , Daohui Wang , Dayun Lin , Duo Zhao , Fengshao Zou , Fu Wang , Gangqiang Zhang , Gengyuan Dan , Guanjie Chen , Guodong Guan , Guodong Yang , Haifeng Li , Haipei Zhu , Haley Li , Hao Feng , Hao Huang , Hao Xu , Hengrui Ma , Hengtao Fan , Hui Liu , Jia Li , Jiang Liu , Jiang Xu , Jie Meng , Jinhan Xin , Junhao Hu , Juwei Chen , Lan Yu , Lanxin Miao , Liang Liu , Linan Jing , Lu Zhou , Meina Han , Mingkun Deng , Mingyu Deng , Naitian Deng , Nizhong Lin , Peihan Zhao , Peng Pan , Pengfei Shen , Ping Li , Qi Zhang , Qian Wang , Qin ZhC Qingrong Xia , Qingyi Zhang , Qunchao Fu , Ren Guo , Ruimin Gao , Shaochun Li , Sheng Long , Shentian Li , Shining Wan , Shuai Shen , Shuangfu Zeng , Shuming Jing , Siqi Yang , Song Zhang , Tao Xu , Tianlin Du , Ting Chen , Wanxu Wu , Wei Jiang , Weinan Tong , Weiwei Chen , Wen Peng , Wenli Zhou , Wenquan Yang , Wenxin Liang , Xiang Liu , Xiaoli Zhou , Xin Jin , Xinyu Duan , Xu Li , Xu Zhang , Xusheng Chen , Yalong Shan , Yang Gan , Yao Lu , Yi Deng , Yi Zheng , Ying Xiong , Yingfei Zheng , Yiyun Zheng , Yizhou Shan , Yong Gao , Yong Zhang , Yongqiang Yang , Yuanjin Gong , Yue Yu , Yuetao Chen , Yukun Zhu , Yulong He , Yusu Zhao , Yuyan Wu , Zenan Zhang , Zhaojin Zhuo , Zhaoyang Ji , Zhefeng Wang , Zheng Wang , Zhenan Fan , Zhenhua Yang , Zhenli Sheng , Zhibin Yu , Zhigang Ji , Zhihao Ren , Zhipeng Bian , Zhixia Liu , Zhiyu Dong , Zhonghua Li , Zhou Yu , Zhuoming Shen , Zhuwei Peng , Zi Ye , Zihao Xiang , Zimin Fu , Zixuan Zhang

With the rapid development of cloud computing systems and the increasing complexity of their infrastructure, intelligent mechanisms to detect and mitigate failures in real time are becoming increasingly important. Traditional methods of…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-20 Cheng Ji , Huaiying Luo

Large multimodal models (LMMs) typically employ an encoding module to transform multimodal data inputs into embeddings, which are then fed to language models for further processing. However, efficiently serving LMMs remains highly…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-30 Tianyu Guo , Tianming Xu , Xianjie Chen , Junru Chen , Nong Xiao , Xianwei Zhang