English
Related papers

Related papers: CAFT: Congestion-Aware Fault-Tolerant Load Balanci…

200 papers

We present novel oblivious routing algorithms for both splittable and unsplittable multicommodity flow. Our algorithm for minimizing congestion for \emph{unsplittable} multicommodity flow is the first oblivious routing algorithm for this…

Data Structures and Algorithms · Computer Science 2017-12-07 Michael Schapira , Gal Shahaf

In this work, we look at the impact of router configurations on DCTCP and Cubic traffic when both algorithms share router buffers in the data center. Modern data centers host traffic with mixed congestion controls, including DCTCP and Cubic…

Networking and Internet Architecture · Computer Science 2023-02-14 Santiago Vargas , Aruna Balasubramanian , Srikanth Sundaresan

Recent years have seen a great increase in the capacity and parallel processing power of data centers and cloud services. To fully utilize the said distributed systems, optimal load balancing for parallel queuing architectures must be…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-08-10 Anam Tahir , Kai Cui , Heinz Koeppl

The increasing penetration of electric vehicles over the coming decades, taken together with the high cost to upgrade local distribution networks and consumer demand for home charging, suggest that managing congestion on low voltage…

Optimization and Control · Mathematics 2015-09-30 Rui Carvalho , Lubos Buzna , Richard Gibbens , Frank Kelly

The exponential growth in model sizes has significantly increased the communication burden in Federated Learning (FL). Existing methods to alleviate this burden by transmitting compressed gradients often face high compression errors, which…

Machine Learning · Computer Science 2025-02-06 Yuhao Zhou , Yuxin Tian , Mingjia Shi , Yuanxi Li , Yanan Sun , Qing Ye , Jiancheng Lv

The Low Latency Fault Tolerance (LLFT) system provides fault tolerance for distributed applications, using the leader-follower replication technique. The LLFT system provides application-transparent replication, with strong replica…

Distributed, Parallel, and Cluster Computing · Computer Science 2010-08-09 Wenbing Zhao , P. M. Melliar-Smith , L. E. Moser

In safety-critical autonomous systems, data freshness presents a fundamental design challenge. While the Logical Execution Time (LET) paradigm ensures compositional determinism, it often does so at the cost of injected latency, degrading…

Operating Systems · Computer Science 2026-03-11 José Luis Conradi Hoffmann , Antônio Augusto Fröhlich

Today's datacenters face important challenges for providing low-latency high-quality interactive services to meet user's expectation. For improving the application throughput, recent research works have embedded application deadline…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-02-04 Ming-Hung Chen , Shi-Chen Wang , Cheng-Fu Chou

Although delay-based congestion control protocols such as FAST promise to deliver better performance than traditional TCP Reno, they have not yet been widely incorporated to the Internet. Several factors have contributed to their lack of…

Networking and Internet Architecture · Computer Science 2014-05-22 Miguel Rodríguez Pérez , Sergio Herrería-Alonso , Manuel Fernández-Veiga , Cándido López-García

A major concern in modern power systems is that the popularity and fluctuating characteristics of renewable energy may cause more and more transmission congestion events. Traditional congestion management modeling involves AC or DC power…

Systems and Control · Electrical Eng. & Systems 2022-12-27 Shutong Pu , Guangchun Ruan , Xinfei Yan , Haiwang Zhong

The behavior of loss-based TCP congestion control algorithms like TCP CUBIC continues to be a challenge in modern cellular networks. Due to the large RLC layer buffers required to deal with short-term changes in channel capacity, the…

Networking and Internet Architecture · Computer Science 2025-10-30 Lukas Prause , Mark Akselrod

Efficient and accurate path-sensitive analyses pose the challenges of: (a) analyzing an exponentially-increasing number of paths in a control-flow graph (CFG), and (b) checking feasibility of paths in a CFG. We address these challenges by…

Software Engineering · Computer Science 2015-03-10 Ahmed Tamrawi , Suresh Kothari

Coflow is a prominent network abstraction for modeling communication patterns in data centers. Since coflow scheduling in large-scale data centers is $\mathcal{NP}$-hard, this paper investigates this problem within heterogeneous parallel…

Data Structures and Algorithms · Computer Science 2026-05-26 Chi-Yeh Chen

In the operation of networked control systems, where multiple processes share a resource-limited and time-varying cost-sensitive network, communication delay is inevitable and primarily influenced by, first, the control systems deploying…

Systems and Control · Electrical Eng. & Systems 2020-07-16 Mohammad H. Mamduhi , Dipankar Maity , Sandra Hirche , John S. Baras , Karl H. Johansson

Fair queuing is becoming increasingly prevalent in the internet and has been shown to improve performance in many circumstances. Performance could be improved even more if endpoints could detect the presence of fair queuing on a certain…

Networking and Internet Architecture · Computer Science 2023-06-12 Maximilian Bachl

Reducing the average memory access time is crucial for improving the performance of applications running on multi-core architectures. With workload consolidation this becomes increasingly challenging due to shared resource contention.…

Hardware Architecture · Computer Science 2021-02-24 Nadja Ramhöj Holtryd , Madhavan Manivannan , Per Stenström , Miquel Pericàs

We present ASYMP, a distributed graph processing system developed for the timely analysis of graphs with trillions of edges. ASYMP has several distinguishing features including a robust fault tolerance mechanism, a lockless architecture…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-12-29 Eduardo Fleury , Silvio Lattanzi , Vahab Mirrokni , Bryan Perozzi

The increasing complexity of AI workloads, especially distributed Large Language Model (LLM) training, places significant strain on the networking infrastructure of parallel data centers and supercomputing systems. While Equal-Cost Multi-…

Networking and Internet Architecture · Computer Science 2024-10-25 Hasibul Jamil , Abdul Alim , Laurent Schares , Pavlos Maniotis , Liran Schour , Ali Sydney , Abdullah Kayi , Tevfik Kosar , Bengi Karacali

Most of the prior work in massively parallel data processing assumes homogeneity, i.e., every computing unit has the same computational capability, and can communicate with every other unit with the same latency and bandwidth. However, this…

Databases · Computer Science 2020-09-25 Xiao Hu , Paraschos Koutris , Spyros Blanas

This paper presents ExeGPT, a distributed system designed for constraint-aware LLM inference. ExeGPT finds and runs with an optimal execution schedule to maximize inference throughput while satisfying a given latency constraint. By…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-12 Hyungjun Oh , Kihong Kim , Jaemin Kim , Sungkyun Kim , Junyeol Lee , Du-seong Chang , Jiwon Seo