English
Related papers

Related papers: FTI-TMR: A Fault Tolerance and Isolation Algorithm…

200 papers

We conducted the feasibility analysis of utilizing a highly available multi-stage architecture for TSN switches used for sending high priority, mission-critical traffic within a bounded latency instead of traditional single-stage…

Networking and Internet Architecture · Computer Science 2022-09-26 Adnan Ghaderi , Rahul Nandkumar Gore

Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster-wide outages. Existing checkpointing systems rely on monolithic, single-tier storage…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-19 Shujie Han , Feng Jiang , Patrick P. C. Lee , Xiao Zhang , Zhijie Huang , Nannan Zhao , Xiaonan Zhao , Lichen Pan

This paper presents a novel fleet management strategy for battery-powered robot fleets tasked with intra-factory logistics in an autonomous manufacturing facility. In this environment, repetitive material handling operations are subject to…

Robotics · Computer Science 2024-09-10 Mithun Goutham , Stephanie Stockar

Enterprise network traffic typically traverses a sequence of middleboxes forming a service function chain, or simply a chain. Tolerating failures when they occur along chains is imperative to the availability and reliability of enterprise…

Networking and Internet Architecture · Computer Science 2020-02-27 Milad Ghaznavi , Elaheh Jalalpour , Bernard Wong , Raouf Boutaba , Ali Jose Mashtizadeh

Stiff dynamical systems present a challenge for machine-learning reduced-order models (ML-ROMs), as explicit time integration becomes unstable in stiff regimes while implicit integration within learning loops is computationally expensive…

Machine Learning · Computer Science 2026-03-19 Joe Standridge , Daniel Livescu , Paul Cizmas

Fault-tolerant distributed applications require mechanisms to recover data lost via a process failure. On modern cluster systems it is typically impractical to request replacement resources after such a failure. Therefore, applications have…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-26 Lukas Hübner , Demian Hespe , Peter Sanders , Alexandros Stamatakis

Time Reversal (TR) has been proposed as a competitive precoding strategy for low-complexity devices, relying on ultra-wideband waveforms. This transmit processing paradigm can address the need for low power and low complexity receivers,…

Information Theory · Computer Science 2023-09-12 Ali Mokh , George C. Alexandropoulos , Mohamed Kamoun , Abdelwaheb Ourir , Arnaud Tourin , Mathias Fink , Julien de Rosny

This paper deals with the fault detection and isolation (FDI) problem for linear structured systems in which the system matrices are given by zero/nonzero/arbitrary pattern matrices. In this paper, we follow a geometric approach to verify…

Optimization and Control · Mathematics 2020-03-04 Jiajia Jia , Harry L. Trentelman , M. Kanat Camlibel

The management of timing constraints in a real-time operating system (RTOS) is usually realized through a global tick counter. This counter acts as the foundational time unit for all tasks in the systems. In order to establish a connection…

Operating Systems · Computer Science 2025-03-04 Kay Heider , Christian Hakert , Kuan-Hsun Chen , Jian-Jia Chen

Ternary large language models (LLMs), which utilize ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision…

Hardware Architecture · Computer Science 2025-06-03 Akul Malhotra , Sumeet Kumar Gupta

We investigate the problem of constructing fault-tolerant bases in matroids. Given a matroid M and a redundancy parameter k, a k-fault-tolerant basis is a minimum-size set of elements such that, even after the removal of any k elements, the…

Data Structures and Algorithms · Computer Science 2025-06-30 Matthias Bentert , Fedor V. Fomin , Petr A. Golovach , Laure Morelle

We consider the problem of active fault-tolerant control in cyber-physical systems composed of strictly passive linear-time invariant dynamic subsystems. We cast the problem as a constrained optimization problem and propose an augmented…

Systems and Control · Electrical Eng. & Systems 2026-02-10 Wasif H. Syed , Juan E. Machado , Johannes Schiffer

Failures in Task-based Parallel Programming (TBPP) can severely degrade performance and result in incomplete or incorrect outcomes. Existing failure-handling approaches, including reactive, proactive, and resilient methods such as retry and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-03-31 Sicheng Zhou , Zhuozhao Li , Valérie Hayot-Sasson , Haochen Pan , Maxime Gonthier , J. Gregory Pauloski , Ryan Chard , Kyle Chard , Ian Foster

A new high-level implementation independent functional fault model for control faults in microprocessors is introduced. The fault model is based on the instruction set, and is specified as a set of data constraints to be satisfied by test…

Hardware Architecture · Computer Science 2019-07-30 Adeboye Stephen Oyeniran , Raimund Ubar , Maksim Jenihhin , Cemil Cem Gursoy , Jaan Raik

An effective way to improve energy efficiency is to throttle hardware resources to meet a certain performance target, specified as a QoS constraint, associated with all applications running on a multicore system. Prior art has proposed…

Hardware Architecture · Computer Science 2019-11-14 Mehrzad Nejat , Madhavan Manivannan , Miquel Pericas , Per Stenstrom

Recent advancements in neutral atom platforms have enabled exploration of early fault-tolerant (FT) architectures for applications with quantum advantage, such as quantum dynamics simulations. An efficient fault-tolerant architecture has…

The realization of fault-tolerant quantum computation hinges on the ability to execute deep quantum circuits while maintaining gate fidelities consistently above error-correction thresholds. Although neutral-atom arrays have recently…

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-27 Rixin Liu , Xingqi Cui , Kaijian Wang , Xinheng Ding , Zirui Liu , Yuke Wang , Jiarong Xing

Multi-robot coordination is crucial for autonomous systems, yet real-world deployments often encounter various failures. These include both temporary and permanent disruptions in sensing and communication, which can significantly degrade…

Robotics · Computer Science 2025-08-11 Peihan Li , Jiazhen Liu , Yuwei Wu , Lifeng Zhou

To cope with the ever increasing threats of dynamic and adaptive persistent attacks, Fault and Intrusion Tolerance (FIT) is being studied at the hardware level to increase critical systems resilience. Based on state-machine replication, FIT…

Cryptography and Security · Computer Science 2023-01-20 Ahmad T Sheikh , Ali Shoker , Paulo Esteves-Verissimo