English
Related papers

Related papers: Predicting Failures in Multi-Tier Distributed Syst…

200 papers

Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely large enough to observe the failures that matter. Jones et al. (2025) address this by…

Machine Learning · Computer Science 2026-05-18 Will Schwarzer , Scott Niekum

For predictive maintenance, we examine one of the largest public datasets for machine failures derived along with their corresponding precursors as error rates, historical part replacements, and sensor inputs. To simplify the time and…

Machine Learning · Computer Science 2018-12-12 David Noever

Proliferation of advanced metering devices with high sampling rates in distribution grids, e.g., micro-phasor measurement units ({\mu}PMU), provides unprecedented potentials for wide-area monitoring and diagnostic applications, e.g.,…

Machine Learning · Computer Science 2017-03-30 Iman Niazazari , Hanif Livani

Real-time detection and mitigation of technical anomalies are critical for large-scale cloud-native services, where even minutes of downtime can result in massive financial losses and diminished user trust. While customer incidents serve as…

Computation and Language · Computer Science 2026-05-25 Jun Wang , Ziyin Zhang , Rui Wang , Hang Yu , Peng Di , Rui Wang

Diffusion Transformers (DiT) have emerged as a widely adopted backbone for high-fidelity image and video generation, yet their iterative denoising process incurs high computational costs. Existing training-free acceleration methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Hanshuai Cui , Zhiqing Tang , Qianli Ma , Zhi Yao , Weijia Jia

Instant payment infrastructures have stringent performance requirements, processing millions of transactions daily with zero-downtime expectations. Traditional monitoring approaches fail to bridge the gap between technical infrastructure…

Machine Learning · Computer Science 2025-10-28 Lorenzo Porcelli

In modern computing environments, users may have multiple systems accessible to them such as local clusters, private clouds, or public clouds. This abundance of choices makes it difficult for users to select the system and configuration for…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-04-05 Amir Nassereldine , Safaa Diab , Mohammed Baydoun , Kenneth Leach , Maxim Alt , Dejan Milojicic , Izzat El Hajj

Large scale data management systems utilize State Machine Replication to provide fault tolerance and to enhance performance. Fault-tolerant protocols are extensively used in the distributed database infrastructure of large enterprises such…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-06-20 Mohammad Javad Amiri , Sujaya Maiyya , Divyakant Agrawal , Amr El Abbadi

Fault-tolerant distributed systems offer high reliability because even if faults in their components occur, they do not exhibit erroneous behavior. Depending on the fault model adopted, hardware and software errors that do not result in a…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-02-19 Rodrigo R. Barbieri , Enrique S. dos Santos , Gustavo M. D. Vieira

Supercomputing systems today often come in the form of large numbers of commodity systems linked together into a computing cluster. These systems, like any distributed system, can have large numbers of independent hardware components…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Michael Treaster

Today, leveraging the enormous modular power, diversity and flexibility of manycore systems-on-a-chip (SoCs) requires careful orchestration of complex resources, a task left to low-level software, e.g. hypervisors. In current architectures,…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-05-11 Inês Pinto Gouveia , Marcus Völp , Paulo Esteves-Verissimo

Understanding material failure is critical for designing stronger and lighter structures by identifying weaknesses that could be mitigated. Existing full-physics numerical simulation techniques involve trade-offs between speed, accuracy,…

Performance unpredictability is a major roadblock towards cloud adoption, and has performance, cost, and revenue ramifications. Predictable performance is even more critical as cloud services transition from monolithic designs to…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-05-06 Yu Gan , Yanqi Zhang , Kelvin Hu , Dailun Cheng , Yuan He , Meghna Pancholi , Christina Delimitrou

Motivated by increasing penetration of distributed generators (DGs) and fast development of micro-phasor measurement units ({\mu}PMUs), this paper proposes a novel graph-based faulted line identification algorithm using a limited number of…

Systems and Control · Electrical Eng. & Systems 2020-02-26 Ying Zhang , Jianhui Wang , Mohammad Khodayar

In this paper, we consider a problem of failure prediction in the context of predictive maintenance applications. We present a new approach for rare failures prediction, based on a general methodology, which takes into account peculiar…

Machine Learning · Computer Science 2019-05-29 Evgeny Burnaev

Root cause analysis in modern cloud infrastructure demands sophisticated understanding of heterogeneous data sources, particularly time-series performance metrics that involve core failure signatures. While large language models demonstrate…

Artificial Intelligence · Computer Science 2026-01-09 Gijun Park

The need for control strategies that can address dynamic system uncertainty is becoming increasingly important. In this work, we propose a Model Predictive Control by quantifying the risk of failure in our system model. The proposed control…

Systems and Control · Electrical Eng. & Systems 2023-02-17 Mostafa Tavakkoli Anbarani , Efe C. Balta , Rômulo Meira-Góes , Ilya Kovalenko

This paper proposes a time-domain fault location identification method for mixed overhead-underground power distribution systems that can handle challenging fault scenarios such as sub-cycle faults, arcing faults and evolving faults. The…

Systems and Control · Electrical Eng. & Systems 2025-10-23 Ali Shakeri Kahnamouei , Saeed Lotfifard

Cloud computing creates new possibilities for control applications by offering powerful computation and storage capabilities. In this paper, we propose a novel cloud-assisted model predictive control (MPC) framework in which we…

Systems and Control · Electrical Eng. & Systems 2021-06-22 Nan Li , Kaixiang Zhang , Zhaojian Li , Vaibhav Srivastava , Xiang Yin

Failure attribution in LLM-based multi-agent systems aims to identify the steps that contribute to a failed execution. This task remains difficult because a single execution can contain many agent actions and tool calls, failure evidence…

Software Engineering · Computer Science 2026-05-15 Yang Liu , Hongjiang Feng , Junsong Pu , Zhuangbin Chen
‹ Prev 1 4 5 6 7 8 10 Next ›