English
Related papers

Related papers: Reliability Models for Highly Fault-tolerant Stora…

200 papers

The rapid advancement of text-to-image Diffusion Models has led to their widespread public accessibility. However these models, trained on large internet datasets, can sometimes generate undesirable outputs. To mitigate this, approximate…

Machine Learning · Computer Science 2024-11-05 Andrea Schioppa , Emiel Hoogeboom , Jonathan Heek

Networked systems are susceptible to cascading failures, where the failure of an initial set of nodes propagates through the network, often leading to system-wide failures. In this work, we propose a multiplex flow network model to study…

Systems and Control · Electrical Eng. & Systems 2025-04-02 Orkun İrsoy , Osman Yağan

We consider a network whose links have random capacities and in which a certain target amount of flow must be carried from some source nodes to some destination nodes. Each destination node has a fixed demand that must be satisfied and each…

Computation · Statistics 2018-05-10 Zdravko I. Botev , Pierre L'Ecuyer , Bruno Tuffin

Due to the increasing complexity of Multi-Processor Systems on Chip (MPSoCs), system-level design methodologies have got a lot of attention in recent years. However, the significant gap between the system-level reliability analysis and the…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-05-14 Hananeh Aliee , Liang Chen , Mojtaba Ebrahimi , Michael Glaß , Faramarz Khosravi , Mehdi B. Tahoori

Archiving and systematic backup of large digital data generates a quick demand for multi-peta byte scale storage systems. As drive capacities continue to grow beyond the few terabytes range to address the demands of today's cloud, the…

Information Theory · Computer Science 2018-10-26 Suayb S. Arslan

In this paper, we study network reliability in relation to a periodic time-dependent utility function that reflects the system's functional performance. When an anomaly occurs, the system incurs a loss of utility that depends on the…

Information Theory · Computer Science 2023-01-16 Ali Maatouk , Fadhel Ayed , Shi Biao , Wenjie Li , Harvey Bao , Enrico Zio

Creating resilient machine learning (ML) systems has become necessary to ensure production-ready ML systems that acquire user confidence seamlessly. The quality of the input data and the model highly influence the successful end-to-end…

Artificial Intelligence · Computer Science 2023-09-21 Manal Rahal , Bestoun S. Ahmed , Jorgen Samuelsson

Fault tolerance overhead of high performance computing (HPC) applications is becoming critical to the efficient utilization of HPC systems at large scale. HPC applications typically tolerate fail-stop failures by checkpointing. Another…

Distributed, Parallel, and Cluster Computing · Computer Science 2011-06-22 Erlin Yao , Mingyu Chen , Rui Wang , Wenli Zhang , Guangming Tan

Security is one of the most relevant concerns in cloud computing. With the evolution of cyber-security threats, developing innovative techniques to thwart attacks is of utmost importance. One recent method to improve cloud computing…

Cryptography and Security · Computer Science 2019-09-05 Matheus Torquato , Marco Vieira

Python's dynamic nature complicates testing and increases the possibility that some defects evade detection, so an effective fault prediction becomes essential. We examine whether post-release faults can be predicted using modern ML and DL.…

Software Engineering · Computer Science 2026-04-30 Giuseppe De Rosa , Pietro Liguori

The memory consistency model is a fundamental system property characterizing a multiprocessor. The relative merits of strict versus relaxed memory models have been widely debated in terms of their impact on performance, hardware complexity…

Distributed, Parallel, and Cluster Computing · Computer Science 2011-04-07 Alexander Jaffe , Thomas Moscibroda , Laura Effinger-Dean , Luis Ceze , Karin Strauss

Multi-task learning (MTL) in materials science relies on the assumption that physically related properties share learnable representations. We challenge this assumption using a 54,028-sample metal alloy dataset exhibiting extreme task-level…

Machine Learning · Computer Science 2026-02-03 Sungwoo Kang

Current contingency reserve criteria ignore the likelihood of individual contingencies and, thus, their impact on system reliability and risk. This paper develops an iterative approach, inspired by the current security-constrained unit…

Systems and Control · Electrical Eng. & Systems 2022-05-10 Robert Mieth , Yury Dvorkin , Miguel A. Ortega-Vazquez

Extreme weather events and cyberattacks can cause component failures and disrupt the operation of power distribution networks (DNs), during which reconfiguration and load shedding are often adopted for resilience enhancement. This study…

Systems and Control · Electrical Eng. & Systems 2026-03-10 Roshni Anna Jacob , Prithvi Poddar , Jaidev Goel , Souma Chowdhury , Yulia R. Gel , Jie Zhang

This paper proposes an active model-based fault and failure tolerant control scheme for a class of marine vehicles with thruster redundancy. Unlike widely used state and parameter estimation methods, where the estimation errors are utilized…

Systems and Control · Electrical Eng. & Systems 2025-02-03 Ji-Hong Li , Hyungjoo Kang , Min-Gyu Kim , Mun-Jik Lee , Han-Sol Jin , Gun Rae Cho

The central topic of this book is application-level fault-tolerance, that is the methods, architectures, and tools that allow to express a fault-tolerant system in the application software of our computers. Application-level fault-tolerance…

Software Engineering · Computer Science 2016-11-09 Vincenzo De Florio

Embedded systems power many modern applications and must often meet strict reliability, real-time, thermal, and power requirements. Task replication can improve reliability by duplicating a task's execution to handle transient and permanent…

Machine Learning · Computer Science 2025-03-18 Roozbeh Siyadatzadeh , Mohsen Ansari , Muhammad Shafique , Alireza Ejlali

As storage systems grow in size, device failures happen more frequently than ever before. Given the commodity nature of hard drives employed, a storage system needs to tolerate a certain number of disk failures while maintaining data…

Information Theory · Computer Science 2014-05-20 Yan Wang , Xunrui Yin , Xin Wang

The robustness of fault detection algorithms against uncertainty is crucial in the real-world industrial environment. Recently, a new probabilistic design scheme called distributionally robust fault detection (DRFD) has emerged and received…

Optimization and Control · Mathematics 2026-01-16 Yulin Feng , Hailang Jin , Steven X. Ding , Hao Ye , Chao Shang

The rise of transient faults in modern hardware requires system designers to consider errors occurring at runtime. Both hardware- and software-based error handling must be deployed to meet application reliability requirements. The level of…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-08-23 Björn Bönninghoff , Horst Schirmeier