English
Related papers

Related papers: DiSCo: Device-Server Collaborative LLM-Based Text …

200 papers

The modern Internet supports diverse applications with heterogeneous quality of service (QoS) requirements. Low Earth orbit (LEO) satellite constellations offer a promising solution to meet these needs, enhancing coverage in rural areas and…

Networking and Internet Architecture · Computer Science 2025-08-29 Dhiraj Bhattacharjee , Pablo G. Madoery , Abhishek Naik , Halim Yanikomeroglu , Gunes Karabulut Kurt , Stephane Martel , Khaled Ahmed

Nowadays, the efficiency and even the feasibility of traditional load-balancing policies are challenged by the rapid growth of cloud infrastructure and the increasing levels of server heterogeneity. In such heterogeneous systems with many…

Networking and Internet Architecture · Computer Science 2020-03-10 Shay Vargaftik , Isaac Keslassy , Ariel Orda

Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, which separates prompt prefill from auto-regressive decode,…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-11 Yiyuan He , Minxian Xu , Jingfeng Wu , Jianmin Hu , Chong Ma , Min Shen , Le Chen , Chengzhong Xu , Lin Qu , Kejiang Ye

In [1], we proposed a two-tier model for creating and maintaining a distributed hash table (DHT) overlay of smartphones. The purpose was to reduce Internet usage and increase D2D content sharing over LTE-A/5G network. However, the real…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-08-11 Sumit Kumar Tetarave , Somanath Tripathy , R. K. Ghosh

Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and efficient Test-Time Scaling (TTS) methods to unlock their full…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Yi Xin , Siqi Luo , Tianxiang Xu , Qi Qin , Haoxing Chen , Kaiwen Zhu , Zhiwei Zhang , Yangfan He , Rongchao Zhang , Jinbin Bai , Shuo Cao , Bin Fu , Junjun He , Yihao Liu , Yuewen Cao , Xiaohong Liu

In service-oriented architectures, accurately predicting the Quality of Service (QoS) is crucial for maintaining reliability and enhancing user satisfaction. However, significant challenges remain due to existing methods always overlooking…

Artificial Intelligence · Computer Science 2024-10-28 Shengxiang Hu , Guobing Zou , Song Yang , Yanglan Gan , Bofeng Zhang , Yixin Chen

In this paper we introduced Modified Sized-based Queue Management as a dropping scheme that aims to fairly prioritize and allocate more service to VoIP traffic over bulk data like FTP as the former one usually has small packet size with…

Networking and Internet Architecture · Computer Science 2012-12-06 Nadim K. M. Madi , Mohamed Othman

Existing LLM serving strategies can be categorized based on whether prefill and decode phases are disaggregated: non-disaggregated (NoDG) or fully disaggregated (FuDG). However, the NoDG strategy leads to strong prefill-decode interference…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-28 Jiangsu Du , Hongbin Zhang , Taosheng Wei , Zhenyi Zheng , Kaiyi Wu , Zhiguang Chen , Yutong Lu

The rapid expansion of web content has made on-device AI assistants indispensable for helping users manage the increasing complexity of online tasks. The emergent reasoning ability in large language models offer a promising path for…

Computation and Language · Computer Science 2025-02-10 Chenyang Shao , Xinyuan Hu , Yutang Lin , Fengli Xu

Multicast short video streaming can enhance bandwidth utilization by enabling simultaneous video transmission to multiple users over shared wireless channels. The existing network management schemes mainly rely on the sequential buffering…

Networking and Internet Architecture · Computer Science 2024-01-24 Xinyu Huang , Shisheng Hu , Haojun Yang , Xinghan Wang , Yingying Pei , Xuemin Shen

- The aim of this paper is to propose a novel Voice On Demand (VoD) architecture and implementation of an efficient load sharing algorithm to achieve Quality of Service (QoS). This scheme reduces the transmission cost from the Centralized…

Multimedia · Computer Science 2016-09-08 M. Dakshayini , T. R. Gopalakrishan Nair

Large Language Models (LLMs) are key technologies driving intelligent systems to handle multiple tasks. To meet the demands of various tasks, an increasing number of LLMs-driven experts with diverse capabilities have been developed,…

Artificial Intelligence · Computer Science 2024-12-06 Yuanshuai Wang , Xingjian Zhang , Jinkun Zhao , Siwei Wen , Peilin Feng , Shuhao Liao , Lei Huang , Wenjun Wu

The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer immense computational power, they incur high capital and…

Artificial Intelligence · Computer Science 2025-10-01 Jovan Stojkovic , Chaojie Zhang , Íñigo Goiri , Ricardo Bianchini

Realistic evaluation of LLM serving systems requires online workloads, dynamic arrivals, queueing, and the serving engine's local scheduling for execution batching, but running such experiments on GPUs is expensive. Existing simulators…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-04 Wei Da , Evangelia Kalyvianaki

Efficient LLM serving must balance throughput and latency across diverse, bursty workloads. We introduce StreamServe, a disaggregated prefill decode serving architecture that combines metric aware routing across compute lanes with adaptive…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-14 Satyam Kumar , Arpit Singh Gautam , Kailash Talreja , Saurabh Jha

Large Language Models (LLMs) have emerged as a transformative force in artificial intelligence, demonstrating exceptional proficiency across various tasks. However, their deployment in resource-constrained environments and concerns over…

Computation and Language · Computer Science 2025-11-11 Tao Fan , Weijing Chen , Yan Kang , Guoqiang Ma , Hanlin Gu , Yuanfeng Song , Lixin Fan , Qiang Yang

Large language models (LLMs) have become essential for applications such as text summarization, sentiment analysis, and automated question-answering. Recently, LLMs have also been integrated into relational database management systems to…

Databases · Computer Science 2025-08-29 Kerem Akillioglu , Anurag Chakraborty , Sairaj Voruganti , M. Tamer Özsu

Recent advances in long-text understanding have pushed the context length of large language models (LLMs) up to one million tokens. It boosts LLMs's accuracy and reasoning capacity but causes exorbitant computational costs and…

Computation and Language · Computer Science 2025-05-19 Huan Yang , Renji Zhang , Mingzhe Huang , Weijun Wang , Yin Tang , Yuanchun Li , Yunxin Liu , Deyu Zhang

LTE evolved Multimedia Broadcast/Multicast Service (eMBMS) is an attractive solution for video delivery to very large groups in crowded venues. However, deployment and management of eMBMS systems is challenging, due to the lack of realtime…

Networking and Internet Architecture · Computer Science 2017-01-12 Yigal Bejerano , Chandru Raman , Chun-Nam Yu , Varun Gupta , Craig Gutterman , Tomas Young , Hugo Infante , Yousef Abdelmalek , Gil Zussman

Text Style Transfer (TST) seeks to alter the style of text while retaining its core content. Given the constraints of limited parallel datasets for TST, we propose CoTeX, a framework that leverages large language models (LLMs) alongside…

Computation and Language · Computer Science 2024-05-07 Chiyu Zhang , Honglong Cai , Yuezhang , Li , Yuexin Wu , Le Hou , Muhammad Abdul-Mageed