中文
相关论文

相关论文: TwinCG: Dual Thread Redundancy with Forward Recove…

200 篇论文

Triple Modular Redundancy (TMR) has been traditionally used to ensure complete tolerance to a single fault or a faulty processing unit, where the processing unit may be a circuit or a system. However, TMR incurs more than 200% overhead in…

硬件体系结构 · 计算机科学 2023-11-02 P Balasubramanian , D L Maskell

Fault tolerance overhead of high performance computing (HPC) applications is becoming critical to the efficient utilization of HPC systems at large scale. HPC applications typically tolerate fail-stop failures by checkpointing. Another…

分布式、并行与集群计算 · 计算机科学 2011-06-22 Erlin Yao , Mingyu Chen , Rui Wang , Wenli Zhang , Guangming Tan

The advances in IC process make future chip multiprocessors (CMPs) more and more vulnerable to transient faults. To detect transient faults, previous core-level schemes provide redundancy for each core separately. As a result, they may…

硬件体系结构 · 计算机科学 2012-06-12 Lei Li , Tianshi Chen , Yunji Chen , Ling Li , Ruiyang Wu

We propose a parallel adaptive constraint-tightening approach to solve a linear model predictive control problem for discrete-time systems, based on inexact numerical optimization algorithms and operator splitting methods. The underlying…

最优化与控制 · 数学 2015-03-24 Laura Ferranti , Tamas Keviczky

Serial-parallel redundancy is a reliable way to ensure service and systems will be available in cloud computing. That method involves making copies of the same system or program, with only one remaining active. When an error occurs, the…

分布式、并行与集群计算 · 计算机科学 2024-04-08 Gutha Jaya Krishna

This paper presents a multilevel convergence framework for multigrid-reduction-in-time (MGRIT) as a generalization of previous two-grid estimates. The framework provides a priori upper bounds on the convergence of MGRIT V- and F-cycles,…

The CMOS integrated chips at advanced technology nodes are becoming more vulnerable to various sources of faults like manufacturing imprecisions, variations, aging, etc. Additionally, the intentional fault attacks (e.g., high power…

硬件体系结构 · 计算机科学 2018-07-08 Naveen Kumar Macha , Bhavana Tejaswini Repalle , Sandeep Geedipally , Rafael Rios , Mostafizur Rahman

Winograd is generally utilized to optimize convolution performance and computational efficiency because of the reduced multiplication operations, but the reliability issues brought by winograd are usually overlooked. In this work, we…

机器学习 · 计算机科学 2023-08-17 Xinghua Xue , Cheng Liu , Bo Liu , Haitong Huang , Ying Wang , Tao Luo , Lei Zhang , Huawei Li , Xiaowei Li

The ever growing demands of embedded systems to satisfy high computing performance and cost efficiency lead to the trend of using commercial off-the-shelf hardware. However, due to their highly integrated design they are becoming…

软件工程 · 计算机科学 2015-11-24 Andrea Höller , Tobias Rauter , Johannes Iber , Georg Macher , Christian Kreiner

We introduce and analyze different strategies for the parallel-in-time integration method PFASST to recover from hard faults and subsequent data loss. Since PFASST stores solutions at multiple time steps on different processors, information…

分布式、并行与集群计算 · 计算机科学 2017-03-21 Robert Speck , Daniel Ruprecht

Parallel-in-time methods for partial differential equations (PDEs) have been the subject of intense development over recent decades, particularly for diffusion-dominated problems. It has been widely reported in the literature, however, that…

数值分析 · 数学 2023-03-22 H. De Sterck , R. D. Falgout , O. A. Krzysik , J. B. Schroder

The two promising methods for capturing high-speed flows are local artificial diffusivity (LAD) and centralised gradient-based reconstruction (C-GBR), the former being computationally economical and the latter being more robust and stable…

流体动力学 · 物理学 2025-11-25 R. R. Kumar , S. Saini , N. R. Vadlamani , A. S. Chamarthi

Distribution system integrated community microgrids (CMGs) can partake in restoring loads during extended duration outages. At such times, the CMG is challenged with limited resource availability, absence of robust grid support, and…

系统与控制 · 电气工程与系统科学 2022-02-11 Ashwin Shirsat , Valliappan Muthukaruppan , Rongxing Hu , Victor Paduani , Bei Xu , Lidong Song , Yiyan Li , Ning Lu , Mesut Baran , David Lubkeman , Wenyuan Tang

Hazard radiation can lead the system fault therefore Fault Tolerance is required. Fault Tolerant is a system, which is designed to keep operations running, despite the degradation in the specific module is happening. Many fault tolerances…

硬件体系结构 · 计算机科学 2014-03-11 Haryono , Jazi Eko Istiyanto , Agus Harjoko , Agfianto Eko Putra

Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in large-scale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for…

分布式、并行与集群计算 · 计算机科学 2025-08-14 Huangliang Dai , Shixun Wu , Jiajun Huang , Zizhe Jian , Yue Zhu , Haiyang Hu , Zizhong Chen

Several recent papers have introduced a periodic verification mechanism to detect silent errors in iterative solvers. Chen [PPoPP'13, pp. 167--176] has shown how to combine such a verification mechanism (a stability test checking the…

数据结构与算法 · 计算机科学 2015-11-17 Massimiliano Fasi , Julien Langou , Yves Robert , Bora Ucar

The increase in HPC systems size and complexity, together with increasing on-chip transistor density, power limitations, and number of components, render modern HPC systems subject to soft errors. Silent data corruptions (SDCs) are…

分布式、并行与集群计算 · 计算机科学 2019-09-04 Aurélien Cavelan , Florina M. Ciorba

The increasing size of deep learning models has made distributed training across multiple devices essential. However, current methods such as distributed data-parallel training suffer from large communication and synchronization overheads…

机器学习 · 计算机科学 2025-02-10 Cabrel Teguemne Fokam , Khaleelulla Khan Nazeer , Lukas König , David Kappel , Anand Subramoney

As computers reach exascale and beyond, the incidence of faults will increase. Solutions to this problem are an active research topic. We focus on strategies to make the preconditioned conjugate gradient (PCG) solver resilient against node…

分布式、并行与集群计算 · 计算机科学 2020-07-09 Carlos Pachajoa , Christina Pacher , Markus Levonyak , Wilfried N. Gansterer

In multi-task learning (MTL), gradient conflict poses a significant challenge. Effective methods for addressing this problem, including PCGrad, CAGrad, and GradNorm, in their original implementations are computationally demanding, which…

机器学习 · 计算机科学 2026-04-03 Evgeny Alves Limarenko , Anastasiia Studenikina , Svetlana Illarionova , Maxim Sharaev