English
Related papers

Related papers: Reliability Models for Highly Fault-tolerant Stora…

200 papers

In future 6G networks, dependable networks will enable telecommunication services such as remote control of robots or vehicles with strict requirements on end-to-end network performance in terms of delay, delay variation, tail…

Networking and Internet Architecture · Computer Science 2026-02-18 Xiaoyu Lan , Jalil Taghia , Hannes Larsson , Andreas Johnsson

Understanding material failure is critical for designing stronger and lighter structures by identifying weaknesses that could be mitigated. Existing full-physics numerical simulation techniques involve trade-offs between speed, accuracy,…

Self-admitted technical debt (SATD), referring to comments flagged by developers that explicitly acknowledge suboptimal code or incomplete functionality, has received extensive attention in machine learning (ML) and traditional (Non-ML)…

Software Engineering · Computer Science 2026-01-21 Niruthiha Selvanayagam , Taher A. Ghaleb , Manel Abdellatif

Foundation model reliability assessment typically requires thousands of evaluation examples, making it computationally expensive and time-consuming for real-world deployment. We introduce microprobe, a novel approach that achieves…

Artificial Intelligence · Computer Science 2025-12-25 Aayam Bansal , Ishaan Gangwani

We propose a novel capacity model for complex networks against cascading failure. In this model, vertices with both higher loads and larger degrees should be paid more extra capacities, i.e. the allocation of extra capacity on vertex $i$…

Physics and Society · Physics 2008-04-03 Ping Li , Bing-Hong Wang , Han Sun , Pan Gao , Tao Zhou

This paper summarizes our work on characterizing application memory error vulnerability to optimize datacenter cost via Heterogeneous-Reliability Memory (HRM), which was published in DSN 2014, and examines the work's significance and future…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-05-11 Yixin Luo , Sriram Govindan , Bikash Sharma , Mark Santaniello , Justin Meza , Aman Kansal , Jie Liu , Badriddine Khessib , Kushagra Vaid , Onur Mutlu

Reliability in multi-agent systems (MAS) built on large language models is increasingly limited by cognitive failures rather than infrastructure faults. Existing observability tools describe failures but do not quantify how quickly…

Multiagent Systems · Computer Science 2025-12-30 Barak Or

We formulate intrusion tolerance for a system with service replicas as a two-level optimal control problem. On the local level node controllers perform intrusion recovery, and on the global level a system controller manages the replication…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-06-06 Kim Hammar , Rolf Stadler

Operating systems include many heuristic algorithms designed to improve overall storage performance and throughput. Because such heuristics cannot work well for all conditions and workloads, system designers resorted to exposing numerous…

A common practice of ML systems development concerns the training of the same model under different data sets, and the use of the same (training and test) sets for different learning models. The first case is a desirable practice for…

Logic in Computer Science · Computer Science 2025-06-06 Leonardo Ceragioli , Giuseppe Primiero

Computing the reliability of a time-varying network, taking into account its dynamic nature, is crucial for networks that change over time, such as space networks, vehicular ad-hoc networks, and drone networks. These networks are modeled…

Data Structures and Algorithms · Computer Science 2025-04-03 Yu Nakahata , Shun Arizono , Shoji Kasahara

Reliability estimation of Machine Learning (ML) models is becoming a crucial subject. This is particularly the case when such \mbox{models} are deployed in safety-critical applications, as the decisions based on model predictions can result…

Machine Learning · Computer Science 2022-06-23 Mohammed Naveed Akram , Akshatha Ambekar , Ioannis Sorokos , Koorosh Aslansefat , Daniel Schneider

The fidelity and utility of synthetic network traffic are critically compromised by architectural mismatch across heterogeneous network datasets and prevalent scalability failure. This study addresses this challenge by establishing an…

Cryptography and Security · Computer Science 2026-03-16 Dure Adan Ammara , Jianguo Ding , Kurt Tutschku

One of the important problem in reliability analysis is computation of stress-strength reliability. But it is impractical to compute it in certain situations. So the estimation stay as an alternative solution to get an approximate value of…

Methodology · Statistics 2022-12-16 Beenu Thomas , V. M. Chacko

In this paper, we examine the different measures of Fault Tolerance in a Distributed Simulated Annealing process. Optimization by Simulated Annealing on a distributed system is prone to various sources of failure. We analyse simulated…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-01-01 Aaditya Prakash

Cold standby 1-out-of-n redundant systems are well-established models in system reliability engineering. To date, reliability analyses of such systems have predominantly assumed exponential, Erlang, or Weibull failure distributions for…

Applications · Statistics 2025-12-30 Afshin Yaghoubi , Esmaile Khorram , Omid Naghshineh Arjmand

The routing algorithms for parallel computers, on-chip networks, multi-core processors, and multiprocessors system-on-chip (MP-SoCs) exhibit router failures must be able to handle interconnect router failures that render a symmetrical mesh…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-01-28 Farshad Safaei , Majed ValadBeigi

Geo-distributed systems often replicate data at multiple locations to achieve availability and performance despite network partitions. These systems must accept updates at any replica and propagate these updates asynchronously to every…

Programming Languages · Computer Science 2019-03-18 Constantin Enea , Suha Orhun Mutluergil , Gustavo Petri , Chao Wang

Fault tolerance is essential for building reliable services; however, it comes at the price of redundancy, mainly the "replication factor" and "diversity". With the increasing reliance on Internet-based services, more machines (mainly…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-04 Ali Shoker

Model adaptation to production environment is critical for reliable Machine Learning Operations (MLOps), less attention is paid to developing systematic framework for updating the ML models when they fail under data drift. This paper…

Machine Learning · Computer Science 2026-02-03 Waqar Muhammad Ashraf , Talha Ansar , Fahad Ahmed , Jawad Hussain , Muhammad Mujtaba Abbas , Vivek Dua