Related papers: Exact Failure Frequency Calculations for Extended …
Fault tolerance overhead of high performance computing (HPC) applications is becoming critical to the efficient utilization of HPC systems at large scale. HPC applications typically tolerate fail-stop failures by checkpointing. Another…
Writing formal specifications for distributed systems is difficult. Even simple consistency requirements often turn out to be unrealizable because of the complicated information flow in the distributed system: not all information is…
This paper presents a unified analysis of the mixed radio-frequency (RF)/free-space optics (FSO) relaying system, with multiple variable-gain amplify-and-forward relays. The partial relay selection (PRS) is employed to select the active…
Most theoretical analysis for lifetime distribution explains origins of specific distribution based on independent failure. We develop a unified framework encompassing different lifetime distribution for failure-coupled network systems. We…
Accurately predicting machine failures in advance can decrease maintenance cost and help allocate maintenance resources more efficiently. Logistic regression was applied to predict machine state 24 hours in the future given the current…
Fault Tree Analysis (FTA) is a well-established method in failure analysis and is widely used in safety and reliability assessments. While FTA tools enable users to manage complex analyses effectively, they can sometimes obscure the…
Failure rates in high performance computers rapidly increase due to the growth in system size and complexity. Hence, failures became the norm rather than the exception. Different approaches on high performance computing (HPC) systems have…
For large-scale power networks, the failure of particular transmission lines can offload power to other lines and cause self-protection trips to activate, instigating a cascade of line failures. In extreme cases, this can bring down the…
In this paper, we propose a price staleness factor model that accounts for pervasive market friction across assets and incorporates relevant covariates. Using large-panel high-frequency data, we derive the maximum likelihood estimators of…
The Meantime to Failure is a statistic used to determine how much time a system spends to enter one of its absorption states. This statistic can be used in most areas of knowledge. In engineering, for example, can be used as a measure of…
Redundancy is widely used to sustain service continuity in programmable and virtualized networks; however, replicated functions often share platforms, software stacks, and control dependencies, making them vulnerable to correlated failures.…
Empirical estimation of critical points at which complex systems abruptly flip from one state to another is among the remaining challenges in network science. However, due to the stochastic nature of critical transitions it is widely…
Engineering networks fall into the category of large-scale networks with heterogeneous nodes such as sources and sinks. The survivability analysis of such networks requires the analysis of the connectivity of the network components for…
We study the reliability of phase oscillator networks in response to fluctuating inputs. Reliability means that an input elicits essentially identical responses upon repeated presentations, regardless of the network's initial condition. In…
We propose a set of dependence measures that are non-linear, local, invariant to a wide range of transformations on the marginals, can show tail and risk asymmetries, are always well-defined, are easy to estimate and can be used on any…
The paper discusses the relationships between electrical and affine differential geometry quantities, establishing a link between frequency and time derivatives of voltage, through the utilization of affine geometric invariants. Based on…
A theorem that describes the high signal-to-noise ratio (SNR) outage behavior of fixed-gain amplify-and-forward (FGAF) relay systems is given. Qualitatively, the theorem states that the outage probability decays according to a power law,…
This work proposes a new and flexible unreliable failure detector whose output is related to the trust level of a set of processes. By expressing the relevance of each process of the set by an impact factor value, our approach allows the…
We present a cascading failure model of two interdependent networks in which functional nodes belong to components of size greater than or equal to $s$. We find theoretically and via simulation that in complex networks with random…
The selective frequency damping method was applied to a bent flow. The method was used in an adaptive formulation. The most dangerous frequency was determined by solving an eigenvalue problem. It was found that one of the patterns,…