English
Related papers

Related papers: Design and Simulation of Fault-Tolerant Network Sw…

200 papers

We will present a new general framework for robust and adaptive control that allows for distributed and scalable learning and control of large systems of interconnected linear subsystems. The control method is demonstrated for a linear…

Systems and Control · Computer Science 2019-04-02 Dimitar Ho , John C. Doyle

By extending the traditional store-and-forward mechanism, network coding has the capability to improve a network's throughput, robustness, and security. Given the fundamentally different packet processing required by this new paradigm and…

Networking and Internet Architecture · Computer Science 2019-09-06 Diogo Gonçalves , Salvatore Signorello , Fernando M. V. Ramos , Muriel Médard

Data streams in real-world industrial scenarios often contain transitional operating conditions that are uncovered during offline training, leading to significant distribution shifts. To bridge the gap between static offline models and…

Systems and Control · Electrical Eng. & Systems 2026-05-26 Hongshuo Zhao , Zeyi Liu , Xiao He

Scientific workflows have been predominantly used for complex and large scale data analysis and scientific computation/automation and the need for robust workflow scheduling techniques has grown considerably. But, most of the existing…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-11-04 S. Jaya Nirmala , Amrith Rajagopal Setlur , Har Simrat Singh , Sudhanshu Khoriya

To support N-1 pre-fault transient stability assessment, this paper introduces a new data collection method in a data-driven algorithm incorporating the knowledge of power system dynamics. The domain knowledge on how the disturbance effect…

Systems and Control · Electrical Eng. & Systems 2022-03-08 Seyedali Meghdadi , Guido Tack , Ariel Liebman , Nicolas Langrené , Christoph Bergmeir

In large-scale LLM pre-training systems with 100k+ GPUs, failures become the norm rather than the exception, and restart costs can dominate wall-clock training time. However, existing fault-tolerance mechanisms are largely unprepared for…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-29 Jin Lee , Zhonghao Chen , Xuhang He , Robert Underwood , Bogdan Nicolae , Franck Cappello , Xiaoyi Lu , Sheng Di , Zheng Zhang

The development of next-generation networks is revolutionizing network operators' management and orchestration practices worldwide. The critical services supported by these networks require increasingly stringent performance requirements,…

Networking and Internet Architecture · Computer Science 2024-05-24 Dimitrios Michael Manias , Joe Naoum-Sawaya , Abbas Javadtalab , Abdallah Shami

We investigate the stability problem for discrete-time stochastic switched linear systems under the specific scenarios where information about the switching patterns and the probability of switches are not available. Our analysis focuses on…

Systems and Control · Computer Science 2018-04-23 Ahmet Cetinkaya , Hideaki Ishii , Tomohisa Hayakawa

In this paper, we study fault-tolerant distributed consensus in wireless systems. In more detail, we produce two new randomized algorithms that solve this problem in the abstract MAC layer model, which captures the basic interface and…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-10-09 Calvin Newport , Peter Robinson

Time-triggered switched networks are a deterministic communication infrastructure used by real-time distributed embedded systems. Due to the criticality of the applications running over them, developers need to ensure that end-to-end…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-01-16 Guy Avni , Shubham Goel , Thomas A. Henzinger , Guillermo Rodriguez-Navas

After decades of research, cascading blackouts remain one of the unresolved challenges in the bulk power system operations. A new perspective for measuring the susceptibility of the system to cascading failures is clearly needed. The newly…

Signal Processing · Electrical Eng. & Systems 2021-01-06 Sayed Abdullah Sadat , Mostafa Sahraei-Ardakani

Many-core systems require inter-core communication, and network-on-chips (NoCs) have been demonstrated to provide good scalability. However, not only the distributed structure but also the link switching on the NoCs have imposed a great…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-22 Niklas Ueter , Georg von der Brueggen , Jian-Jia Chen , Tulika Mitra , Vanchinathan Venkataramani

Emergency communications networks require in-network intelligence for timely traffic handling under dynamic demands and runtime constraints. In these environments, packets may need different inference behaviors, and conventional model…

Networking and Internet Architecture · Computer Science 2026-05-12 Yuehan Li , Zhiyuan Ren , Tao Zhang , Wenchi Cheng

Advanced integration of logistics systems has been promoted for the sake of competitiveness and sustainability. Such efforts will enable more globally optimal and flexible operations by efficiently utilizing transportation capacity. At the…

Physics and Society · Physics 2022-03-29 Takahiro Ezaki , Naoto Imura , Katsuhiro Nishinari

This paper presents a methodology for model based robust fault diagnosis and a methodology for input design to obtain optimal diagnosis of faults. The proposed algorithm is suitable for real time implementation. Issues of robustness are…

Systems and Control · Computer Science 2020-01-16 Dhruv Khandelwal , Siep Weiland , Amol Khalate

The structures for the expression of fault-tolerance provisions into the application software are the central topic of this dissertation. Structuring techniques provide means to control complexity, the latter being a relevant factor for the…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-11-08 Vincenzo De Florio

Deep neural networks often exhibit poor performance on data that is unlikely under the train-time data distribution, for instance data affected by corruptions. Previous works demonstrate that test-time adaptation to data shift, for instance…

Wormhole routing, the latest switching technique to be utilized by massively parallel computers, enjoys the distinct advantage of a low latency when compared to other switching techniques. This low latency is due to the nearly distance…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-08-31 Denvil Smith

Queuing network control is essential for managing congestion in job-processing systems such as service systems, communication networks, and manufacturing processes. Despite growing interest in applying reinforcement learning (RL)…

Machine Learning · Computer Science 2024-09-06 Ethan Che , Jing Dong , Hongseok Namkoong

Time-Sensitive Networking (TSN) is increasingly adopted in industrial systems to meet strict latency, jitter, and reliability requirements. However, evaluating TSN's fault tolerance under realistic failure conditions remains challenging.…

Networking and Internet Architecture · Computer Science 2025-07-16 Mohamed Seliem , Dirk Pesch , Utz Roedig , Cormac Sreenan
‹ Prev 1 4 5 6 7 8 10 Next ›