English
Related papers

Related papers: Varuna: Enabling Failure-Type Aware RDMA Failover

200 papers

Among all requirements for vehicle-to-everything (V2X) communications, successful delivery of packets with small delay is of the highest significance. Especially, the delivery of a message before a potential accident (i.e. emergency…

Networking and Internet Architecture · Computer Science 2020-05-13 Junwei Zang , Vahid Towhidlou , Mohammad Shikh-Bahaei

Reliable and secure communication is essential for mission-critical aerospace and defence operations involving autonomous platforms such as Unmanned Aerial Vehicles (UAVs), satellites, and ground control systems. In contested or dynamic…

Systems and Control · Electrical Eng. & Systems 2026-04-29 Dhrumil Bhatt , Anakha Kurup

Failures in Task-based Parallel Programming (TBPP) can severely degrade performance and result in incomplete or incorrect outcomes. Existing failure-handling approaches, including reactive, proactive, and resilient methods such as retry and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-03-31 Sicheng Zhou , Zhuozhao Li , Valérie Hayot-Sasson , Haochen Pan , Maxime Gonthier , J. Gregory Pauloski , Ryan Chard , Kyle Chard , Ian Foster

Most modern communication networks include fast rerouting mechanisms, implemented entirely in the data plane, to quickly recover connectivity after link failures. By relying on local failure information only, these data plane mechanisms…

Networking and Internet Architecture · Computer Science 2024-11-27 Gregor Bankhamer , Robert Elsässer , Stefan Schmid

This paper proposes a new reliability algorithm specifically useful when retransmission is either problematic or not possible. In case of multimedia or multicast communications and in the context of the Delay Tolerant Networking (DTN), the…

Networking and Internet Architecture · Computer Science 2008-09-29 Jerome Lacan , Emmanuel Lochin

Combining persistent memory (PM) with RDMA is a promising approach to performant replicated distributed key-value stores (KVSs). However, existing replication approaches do not work well when applied to PM KVSs: 1) Using RPC induces…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-21 Qing Wang , Youyou Lu , Jing Wang , Jiwu Shu

Serverless computing promises enhanced resource efficiency and lower user costs, yet is burdened by a heavyweight, CPU-bound data plane. Prior efforts exploiting shared memory reduce overhead locally but fall short when scaling across…

Networking and Internet Architecture · Computer Science 2025-05-19 Shixiong Qi , Songyu Zhang , K. K. Ramakrishnan , Diman Z. Tootaghaj , Hardik Soni , Puneet Sharma

The performance of large-scale computing systems often critically depends on high-performance communication networks. Dynamically reconfigurable topologies, e.g., based on optical circuit switches, are emerging as an innovative new…

Networking and Internet Architecture · Computer Science 2022-12-29 Vamsi Addanki , Chen Avin , Stefan Schmid

Modern communication networks feature local fast failover mechanisms in the data plane, to swiftly respond to link failures with pre-installed rerouting rules. This paper explores resilient routing meant to tolerate $\leq k$ simultaneous…

Networking and Internet Architecture · Computer Science 2024-10-04 Wenkai Dai , Klaus-Tycho Foerster , Stefan Schmid

Handling faults is a growing concern in HPC. In future exascale systems, it is projected that silent undetected errors will occur several times a day, increasing the occurrence of corrupted results. In this article, we propose SEDAR, which…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-29 Diego Montezanti , Enzo Rucci , Armando De Giusti , Marcelo Naiouf , Dolores Rexachs , Emilio Luque

When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network performance, is highly inefficient, requiring datacenters to…

Machine Learning · Computer Science 2025-10-07 Abhishek Vijaya Kumar , Eric Ding , Arjun Devraj , Darius Bunandar , Rachee Singh

This paper describes two-fold approach towards utilizing Triple Modular Redundancy (TMR) in Wireless Adhoc Network (AdocNet). A distributed checkpointing and recovery protocol is proposed. The protocol eliminates useless checkpoints and…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-11-11 Sarmistha Neogy

Redundant transfer of resources is a critical issue for compromising the performance of mobile Web applications (a.k.a., apps) in terms of data traffic, load time, and even energy consumption. Evidence shows that the current cache…

Software Engineering · Computer Science 2016-12-12 Xuanzhe Liu , Yun Ma , Shuailiang Dong , Yunxin Liu , Tao Xie , Gang Huang , Hong Mei

This report presents results of our endeavor towards developing a failure-recovery variant of a CORBA-based bank server that provides fault tolerance features through message logging and checkpoint logging. In this group of projects, three…

Distributed, Parallel, and Cluster Computing · Computer Science 2009-11-17 Emil Vassev , Que Thu Dung Nguyen , Heng Kuang

Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime reconfiguration suffer from long downtimes and degraded…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-18 ChonLam Lao , Jiaqi Gao , Jiamin Cao , Zhipeng Zhang , Pengcheng Zhang , Jiangfei Duan , Zhilong Zheng , Yu Guan , Yichi Xu , Yong Li , Zhengping Qian , Aditya Akella , Minlan Yu , Ennan Zhai , Dennis Cai , Jingren Zhou

LoRaWAN deployments follow an ad-hoc deployment model that has organically led to overlapping communication networks, sharing the wireless spectrum, and completely unaware of each other. LoRaWAN uses ALOHA-style communication where it is…

Networking and Internet Architecture · Computer Science 2021-11-22 Laksh Bhatia , Po-Yu Chen , Michael Breza , Cong Zhao , Julie A. McCann

In this paper, we conduct systematic measurement studies to show that the high memory bandwidth consumption of modern distributed applications can lead to a significant drop of network throughput and a large increase of tail latency in…

Single node failures represent more than 85% of all node failures in the today's large communication networks such as the Internet. Also, these node failures are usually transient. Consequently, having the routing paths globally recomputed…

Data Structures and Algorithms · Computer Science 2008-10-21 Amit M Bhosle , Teofilo F Gonzalez

Learning-based navigation systems are widely used in autonomous applications, such as robotics, unmanned vehicles and drones. Specialized hardware accelerators have been proposed for high-performance and energy-efficiency for such…

Robotics · Computer Science 2021-11-10 Zishen Wan , Aqeel Anwar , Yu-Shun Hsiao , Tianyu Jia , Vijay Janapa Reddi , Arijit Raychowdhury

Backpressure scheduling and routing, in which packets are preferentially transmitted over links with high queue differentials, offers the promise of throughput-optimal operation for a wide range of communication networks. However, when the…

Networking and Internet Architecture · Computer Science 2016-11-17 Majed Alresaini , Maheswaran Sathiamoorthy , Bhaskar Krishnamachari , Michael J. Neely