Related papers: The system operating time with two different unrel…
Most existing research about complex systems maintenance assumes they consist of the same type of components. However, systems can be assembled with heterogeneous components (for example degrading and non-degrading components) that require…
We present an approach for verifying systems at runtime. Our approach targets distributed systems whose components communicate with monitors over unreliable channels, where messages can be delayed, reordered, or even lost. Furthermore, our…
Multi-stability is a widely observed phenomenon in real complex networked systems, such as technological infrastructures, ecological systems, gene regulation, transportation and more. When a system functions normally but there exists also a…
In the past, we have observed several large blackouts, i.e. loss of power to large areas. It has been noted by several researchers that these large blackouts are a result of a cascade of failures of various components. As a power grid is…
A semicoherent system can be described by its structure function or, equivalently, by a lattice polynomial function expressing the system lifetime in terms of the component lifetimes. In this paper we point out the parallelism between the…
Runtime performance variability at the servers has been a major issue, hindering the predictable and scalable performance in modern distributed systems. Executing requests or jobs redundantly over multiple servers has been shown to be…
Concurrent accesses to databases are typically encapsulated in transactions in order to enable isolation from other concurrent computations and resilience to failures. Modern databases provide transactions with various semantics…
Optimal maintenance policies play an important role in the reliability analysis of repairable systems. This paper examines a two-unit priority standby system with a repair facility, where the priority unit is subject to preventive…
Power grid outages cause huge economical and societal costs. Disruptions in the power distribution grid are responsible for a significant fraction of electric power unavailability to customers. The impact of extreme weather conditions,…
Customer slowdown describes the phenomenon that a customer's service requirement increases with experienced delay. In healthcare settings, there is substantial empirical evidence for slowdown, particularly when a patient's delay exceeds a…
Distributed storage systems are mainly justified due to the limited amount of storage capacity and improving the reliability through distributing data over multiple storage nodes. On the other hand, it may happen the data is stored in…
A balanced investigation into the reliability of wireless smart health devices when it comes to the collection of biometric data under varying network/environmental conditions. Followed by a program implementation to begin introductory…
We consider a downlink periodic wireless communications system where multiple access points (APs) cooperatively transmit packets to a number of devices, e.g. actuators in an industrial control system. Each period consists of two phases: an…
Control systems behavior can be analyzed taking into account a large number of parameters: performances, reliability, availability, security. Each control system presents various security vulnerabilities that affect in lower or higher…
Noise is often considered to be a nuisance. Here we argue that it can be a useful probe of fluctuating two level systems in glasses. It can be used to: (1) shed light on whether the fluctuations are correlated or independent events; (2)…
Control systems can show robustness to many events, like disturbances and model inaccuracies. It is natural to speculate that they are also robust to sporadic deadline misses when implemented as digital tasks on an embedded platform. This…
With increasing complexity of modern-day mobile devices, security of these devices in presence of myriad attacks by an intelligent adversary is becoming a major issue. The vast majority of cell phones still remain unsecured from many…
Continuous availability of HPC systems built from commodity components have become a primary concern as system size grows to thousands of processors. In this paper, we present the analysis of 8-24 months of real failure data collected from…
This study aims to develop a wearable device that collect health data from maintenance personnel and environmental conditions data in order to ensure the safety of the staff in industrial work areas where have different levels of risk…
Failure detectors are oracles that have been introduced to provide processes in asynchronous systems with information about faults. This information can then be used to solve problems otherwise unsolvable in asynchronous systems. A natural…