English
Related papers

Related papers: HEAL: Online Incremental Recovery for Leaderless D…

200 papers

Space Cyber-Physical Systems (S-CPS) such as spacecraft and satellites strongly rely on the reliability of onboard computers to guarantee the success of their missions. Relying solely on radiation-hardened technologies is extremely…

Systems and Control · Electrical Eng. & Systems 2024-01-12 Michael Rogenmoser , Yvan Tortorella , Davide Rossi , Francesco Conti , Luca Benini

Modern database management systems (DBMS) face significant challenges in maintaining performance and availability under dynamic workloads. This paper proposes a novel self-healing framework that integrates Model-Agnostic Meta-Learning…

Databases · Computer Science 2026-05-06 Joydeep Chandra , Prabal Manhas

Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Unfortunately, existing machine learning systems focus…

Databases · Computer Science 2018-12-17 Doris Xin , Stephen Macke , Litian Ma , Jialin Liu , Shuchen Song , Aditya Parameswaran

A significant fraction of software failures in large-scale Internet systems are cured by rebooting, even when the exact failure causes are unknown. However, rebooting can be expensive, causing nontrivial service disruption or downtime even…

Operating Systems · Computer Science 2007-05-23 George Candea , Shinichi Kawamoto , Yuichi Fujiki , Greg Friedman , Armando Fox

Distributed storage architectures are foundational to modern cloud-native infrastructure, yet a critical operational bottleneck persists within disaster recovery (DR) workflows: the dependence on content-based cryptographic hashing for data…

Cryptography and Security · Computer Science 2026-02-27 Prasanna Kumar , Nishank Soni , Gaurang Munje

As LLM deployments scale over more hardware, the probability of a single failure in a system increases significantly, and cloud operators must consider robust countermeasures to handle these inevitable failures. A common recovery approach…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-25 Haley Li , Xinglu Wang , Cong Feng , Chunxu Zuo , Yanan Wang , Hei Lo , Yufei Cui , Bingji Wang , Duo Cui , Shuming Jing , Yizhou Shan , Ying Xiong , Jiannan Wang , Yong Zhang , Zhenan Fan

Ensuring the reliability and resilience of modern web applications remains a critical challenge due to increasing system complexity and dynamic runtime environments. This study proposes a modular self-healing framework based on the…

Software Engineering · Computer Science 2026-05-20 Sales Aribe , Rov Japheth Oracion

Distributed storage systems introduce redundancy to protect data from node failures. After a storage node fails, the lost data should be regenerated at a replacement storage node as soon as possible to maintain the same level of redundancy.…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-06-19 Qingyuan Gong , Jiaqi Wang , Yan Wang , Dongsheng Wei , Jin Wang , Xin Wang

Database search and clustering are fundamental components of many data analytics problems, such as mass spectrometry-driven proteomics. Traditional full clustering and search algorithms suffer from high resource usage and long latencies. We…

Databases · Computer Science 2025-11-25 Md Mizanur Rahaman Nayan , Zheyu Li , Flavio Ponzina , Sumukh Pinge , Tajana Rosing , Azad J. Naeemi

Scaling supercomputers comes with an increase in failure rates due to the increasing number of hardware components. In standard practice, applications are made resilient through checkpointing data and restarting execution after a failure…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-02-16 Giorgis Georgakoudis , Luanzheng Guo , Ignacio Laguna

This paper proposes a novel method to co-optimize distribution system operation and repair crew routing for outage restoration after extreme weather events. A two-stage stochastic mixed integer linear program is developed. The first stage…

Optimization and Control · Mathematics 2018-06-29 Anmar Arif , Shanshan Ma , Zhaoyu Wang , Jianhui Wang , Sarah M. Ryan , Chen Chen

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved…

Reed relay serves as the fundamental component of functional testing, which closely relates to the successful quality inspection of electronics. To provide accurate remaining useful life (RUL) estimation for reed relay, a hybrid deep…

Machine Learning · Computer Science 2022-09-15 Chinthaka Gamanayake , Yan Qin , Chau Yuen , Lahiru Jayasinghe , Dominique-Ea Tan , Jenny Low

Ameliorating the lifetime in heterogeneous wireless sensor network is an important task because the sensor nodes are limited in the resource energy. The best way to improve a WSN lifetime is the clustering based algorithms in which each…

Networking and Internet Architecture · Computer Science 2014-08-20 Mostafa Baghouri , Saad Chakkor , Abderrahmane Hajraoui

The introduction of electronic personal health records (EHR) enables nationwide information exchange and curation among different health care systems. However, the current EHR systems do not provide transparent means for diagnosis support,…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-10-04 Dragi Kimovski , Sasko Ristov , Radu Prodan

We consider cooperative communications with energy harvesting (EH) relays, and develop a distributed power control mechanism for the relaying terminals. Unlike prior art which mainly deal with single-relay systems with saturated traffic…

Networking and Internet Architecture · Computer Science 2018-10-26 Vesal Hakami , Mehdi Dehghan

Deep learning (DL) accelerators are increasingly deployed on edge devices to support fast local inferences. However, they suffer from a new security problem, i.e., being vulnerable to physical access based attacks. An adversary can easily…

Hardware Architecture · Computer Science 2020-08-11 Pengfei Zuo , Yu Hua , Ling Liang , Xinfeng Xie , Xing Hu , Yuan Xie

Self-healing capability is one of the most critical factors for a resilient distribution system, which requires intelligent agents to automatically perform restorative actions online, including network reconfiguration and reactive power…

Systems and Control · Electrical Eng. & Systems 2021-05-11 Yichen Zhang , Feng Qiu , Tianqi Hong , Zhaoyu Wang , Fangxing Li

More than 70% of cloud computing is paid for but sits idle. A large fraction of these idle compute are cheap CPUs with few cores that are not utilized during the less busy hours. This paper aims to enable those CPU cycles to train…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-02-01 Minghao Yan , Nicholas Meisburger , Tharun Medini , Anshumali Shrivastava

Microservice based systems underpin modern distributed computing environments but remain vulnerable to partial failures, cascading timeouts, and inconsistent recovery behavior. Although numerous resilience and recovery patterns have been…

Software Engineering · Computer Science 2026-02-03 Muzeeb Mohammad