English
Related papers

Related papers: SwiftQueue: Optimizing Low-Latency Applications wi…

200 papers

High load latency that results from deep cache hierarchies and relatively slow main memory is an important limiter of single-thread performance. Data prefetch helps reduce this latency by fetching data up the hierarchy before it is…

Hardware Architecture · Computer Science 2021-03-30 Majid Jalili , Mattan Erez

Federated Learning (FL) allows devices to train a global machine learning model without sharing data. In the context of wireless networks, the inherently unreliable nature of the transmission channel introduces delays and errors that…

Networking and Internet Architecture · Computer Science 2024-08-05 Renan R. de Oliveira , Kleber V. Cardoso , Antonio Oliveira-Jr

Low-latency live streaming (LLS) has emerged as a popular web application, with many platforms adopting real-time protocols such as WebRTC to minimize end-to-end latency. However, we observe a counter-intuitive phenomenon: even when the…

Image and Video Processing · Electrical Eng. & Systems 2026-02-11 Liming Liu , Zhidong Jia , Li Jiang , Wei Zhang , Lan Xie , Feng Qian , Leju Yan , Bing Yan , Qiang Ma , Zhou Sha , Wei Yang , Yixuan Ban , Xinggong Zhang

The time-varying feature of wireless channels usually makes the hard delay bound for data transmissions unrealistic to guarantee. In contrast, the statistically-bounded delay with a small violation probability has been widely used for delay…

Information Theory · Computer Science 2011-05-03 Qinghe Du , Yi Huang , Pinyi Ren , Chao Zhang

The increasing number of different, incompatible congestion control algorithms has led to an increased deployment of fair queuing. Fair queuing isolates each network flow and can thus guarantee fairness for each flow even if the flows'…

Networking and Internet Architecture · Computer Science 2021-01-25 Maximilian Bachl , Joachim Fabini , Tanja Zseby

Large language models (LLMs) deliver impressive capabilities but incur substantial inference latency and cost, which hinders their deployment in latency-sensitive and resource-constrained scenarios. Cloud-edge-device collaborative inference…

Artificial Intelligence · Computer Science 2026-03-24 Haoyu Qiao , Hao Zhang , Shanwen Mao , Siyao Cheng , Jie Liu

Length Rate Quotient (LRQ) is the first algorithm of interleaved shaping -- a novel concept proposed to provide per-flow shaping for a flow aggregate without per-flow queuing. This concept has been adopted by Time-Sensitive Networking (TSN)…

Performance · Computer Science 2021-07-13 Yuming Jiang

Effective internet traffic prediction in smaller ISP networks is challenged by limited data availability. This paper explores this issue using transfer learning and data augmentation techniques with two LSTM-based models, LSTMSeq2Seq and…

Machine Learning · Computer Science 2025-09-22 Sajal Saha , Anwar Haque , Greg Sidebottom

Control of multihop Wireless networks in a distributed manner while providing end-to-end delay requirements for different flows, is a challenging problem. Using the notions of Draining Time and Discrete Review from the theory of fluid…

Networking and Internet Architecture · Computer Science 2017-04-20 Ashok Krishnan K. S. , Vinod Sharma

The current development trend of wireless communications aims at coping with the very stringent reliability and latency requirements posed by several emerging Internet of Things (IoT) application scenarios. Since the problem of realizing…

Networking and Internet Architecture · Computer Science 2025-03-03 Federico Librino , Paolo Santi

In edge-cloud speculative decoding (SD), edge devices equipped with small language models (SLMs) generate draft tokens that are verified by large language models (LLMs) in the cloud. A key bottleneck in such systems is the limited…

Signal Processing · Electrical Eng. & Systems 2026-01-13 Guangyi Zhang , Yunlong Cai , Guanding Yu , Petar Popovski , Osvaldo Simeone

A key feature of the packet scheduler in LTE system is that it can allocate resources both in the time and frequency domain. Furthermore, the scheduler is acquainted with channel state information periodically reported by user equipments…

Networking and Internet Architecture · Computer Science 2014-10-01 Mattia Carpin , Andrea Zanella , Jawad Rasool , Kashif Mahmood , Ole Grøndalen , Olav N. Østerbø

Deploying large language models (LLMs) in mobile and edge computing environments is constrained by limited on-device resources, scarce wireless bandwidth, and frequent model evolution. Although edge-cloud collaborative inference with…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-05 Yuchen Li , Rui Kong , Zhonghao Lyu , Qiyang Li , Xinran Chen , Hengyi Cai , Lingyong Yan , Shuaiqiang Wang , Jiashu Zhao , Guangxu Zhu , Linghe Kong , Guihai Chen , Haoyi Xiong , Dawei Yin

We consider the problem of scheduling transmissions over low-latency wireless communication links to control various control systems. Low-latency requirements are critical in developing wireless technology for industrial control and Tactile…

Signal Processing · Electrical Eng. & Systems 2019-10-31 Mark Eisen , Mohammad M. Rashid , Dave Cavalcanti , Alejandro Ribeiro

Many routing protocols have been proposed to handle reliability and real-time routing energy efficiency for wireless sensor networks. In this paper we propose a new routing protocol with QoS based capabilities for WSNs. We used priority…

Networking and Internet Architecture · Computer Science 2018-03-13 Hossein Pourakbar , Ali Ghaffari

Load balancing algorithms play a vital role in enhancing performance in data centers and cloud networks. Due to the massive size of these systems, scalability challenges, and especially the communication overhead associated with load…

Probability · Mathematics 2019-03-07 Mark van der Boor , Sem Borst , Johan van Leeuwaarden

In view of the fact that routing algorithms are network layer entities and the varying performance of any routing algorithm depends on the underlying networks. Localized routing algorithms avoid the problems associated with the maintenance…

Networking and Internet Architecture · Computer Science 2011-03-18 Abdulbaset H. Mohammad

Max weighted queue (MWQ) control policy is a widely used cross-layer control policy that achieves queue stability and a reasonable delay performance. In most of the existing literature, it is assumed that optimal MWQ policy can be obtained…

Systems and Control · Computer Science 2013-11-20 Junting Chen , Vincent K. N. Lau

Speculative decoding accelerates Large Language Model (LLM) inference by employing a small speculative model (SSM) to generate multiple candidate tokens and verify them using the LLM in parallel. This technique has been widely integrated…

Computation and Language · Computer Science 2025-05-26 Ruixiao Li , Fahao Chen , Peng Li

Recurrent Neural Networks (RNNs) have been shown to be valuable for constructing Intrusion Detection Systems (IDSs) for network data. They allow determining if a flow is malicious or not already before it is over, making it possible to take…

Machine Learning · Computer Science 2020-10-16 Maximilian Bachl , Fares Meghdouri , Joachim Fabini , Tanja Zseby