English
Related papers

Related papers: Reliability Analysis of Fault Tolerant Memory Syst…

200 papers

Deep-learning-based recommendation models (DLRMs) are widely deployed to serve personalized content to users. DLRMs are large in size due to their use of large embedding tables, and are trained by distributing the model across the memory of…

Machine Learning · Computer Science 2021-04-06 Kaige Liu , Jack Kosaian , K. V. Rashmi

Inter-channel mis-synchronisation can be a limiting factor to the time resolution of high performance timing detectors with multiple readout channels and independent electronics units. In these systems, time calibration methods employed…

Instrumentation and Detectors · Physics 2026-03-03 S. Abe , H. Alarakia-Charles , I. Alekseev , C. Alt , T. Arai , T. Arihara , S. Arimoto , A. M. Artikov , Y. Awataguchi , N. Babu , V. Baranov , G. Barr , D. Barrow , L. Bartoszek , L. Bernardi , L. Berns , S. Bhattacharjee , A. V. Boikov , A. Blanchet , A. Blondel , A. Bonnemaison , S. Bordoni , M. H. Bui , T. H. Bui , F. Cadoux , S. Cap , A. Cauchois , J. Chakrani , P. S. Chong , A. Chvirova , P. Collard , M. Danilov , C. Davis , V. Davouloury , Yu. I. Davydov , A. Dergacheva , C. Domangue , D. Douqa , T. A. Doyle , O. Drapier , A. Eguchi , J. Elias , G. Erofeev , Y. Favre , D. Fedorova , S. Fedotov , D. Ferlewicz , Y. Fujii , R. Fujita , Y. Furui , F. Gastaldi , A. Gendotti , A. Germer , L. Giannessi , C. Giganti , V. Glagolev , R. Guillaumat , G. Ha , N. C. Hastings , I. Heitkamp , J. Hu , C. Husi , A. K. Ichikawa , T. H. Ishida , A. Izmaylov , K. Iwamoto , M. Jakkapu , C. Jesús-Valls , J. Y. Ji , P. Jonsson , C. K. Jung , H. Kakuno , V. S. Kasturi , M. Kawaue , P. T. Keener , M. Khabibullin , N. V. Khomutov , A. Khotjantsev , T. Kikawa , H. Kikutani , N. V. Kirichkov , A. Klustová , H. Kobayashi , T. Kobayashi , L. Koch , S. Kodama , A. O. Kolesnikov , M. Kolupanova , A. Korzenev , T. Koto , Y. Kudenko , S. Kuribayashi , T. Kutter , M. Lachat , K. Lachner , M. Lamers James , D. Last , N. Latham , M. Lawe , T. A. Le , D. Leon Silverio , B. Li , W. Li , C. Lin , M. Louzir , T. Lux , K. K. Mahtani , S. Manly , D. A. Martinez Caicedo , N. Mashin , T. Matsubara , C. Mauger , K. S. McFarland , C. McGrew , J. McKean , A. Mefodiev , E. Miller , O. Mineev , A. Minamino , A. L. Moreno , A. Muñoz , T. Nakadaira , K. Nakagiri , T. Nakaya , J. Nanni , L. Nicolas , A. D. Nguyen , D. T. Nguyen , H. Nguyen , V. Nguyen , E. Noah Messomo , T. Nosek , H. M. O'Keeffe , T. Ogawa , W. Okinaga , L. Osu , V. Paolone , G. Pelleriti , L. Pickering , M. A. Ramírez , M. Reh , G. Reina , C. Riccio , S. Roth , A. Rubbia , F. Saadi , K. Sakashita , N. Sallin , S. Samani , F. Sanchez , T. Schefke , C. Schloesser , D. Sgalaberna , A. Shaikovskiy , A. Shvartsman , Y. Shiraishi , N. Shvarev , N. Skrobova , D. Smyczek , M. Smy , A. Speers , D. Svirida , M. Ta , S. Tairafune , M. Tani , H. Tanigawa , A. Teklu , S. Tereshchenko , V. V. Tereshchenko , T. Thaiduc , T. Tsushima , M. Tzanov , Y. Uchida , I. I. Vasilyev , E. Villa , T. Vladisavljevic , D. Wakabayashi , H. Wallace , A. Weber , N. Whitney , C. Wret , Y. Xu , Y. Yang , N. Yershov , A. J. P. Yrey , M. Yokoyama , Y. Yoshimoto , X. Y. Zhao , H. Zheng , H. Zhong , T. Zhu , E. D. Zimmerman , M. Zito

An open research question in deep reinforcement learning is how to focus the policy learning of key decisions within a sparse domain. This paper emphasizes combining the advantages of inputoutput hidden Markov models and reinforcement…

Machine Learning · Computer Science 2023-01-12 Ammar N. Abbas , Georgios Chasparis , John D. Kelleher

We introduce a new class of exact Minimum-Bandwidth Regenerating (MBR) codes for distributed storage systems, characterized by a low-complexity uncoded repair process that can tolerate multiple node failures. These codes consist of the…

Information Theory · Computer Science 2010-10-14 Salim El Rouayheb , Kannan Ramchandran

The advances in IC process make future chip multiprocessors (CMPs) more and more vulnerable to transient faults. To detect transient faults, previous core-level schemes provide redundancy for each core separately. As a result, they may…

Hardware Architecture · Computer Science 2012-06-12 Lei Li , Tianshi Chen , Yunji Chen , Ling Li , Ruiyang Wu

The robustness of distributed optimization is an emerging field of study, motivated by various applications of distributed optimization including distributed machine learning, distributed sensing, and swarm robotics. With the rapid…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-06-29 Shuo Liu

We describe verification techniques for embedded memory systems using efficient memory modeling (EMM), without explicitly modeling each memory bit. We extend our previously proposed approach of EMM in Bounded Model Checking (BMC) for a…

Logic in Computer Science · Computer Science 2011-11-09 Malay K. Ganai , Aarti Gupta , Pranav Ashar

Machine unlearning aims to selectively remove targeted knowledge from Large Language Models (LLMs), ensuring they forget specified content while retaining essential information. Existing unlearning metrics assess whether a model correctly…

Computation and Language · Computer Science 2025-05-28 Wonje Jeung , Sangyeon Yoon , Albert No

Classical erasure codes, e.g. Reed-Solomon codes, have been acknowledged as an efficient alternative to plain replication to reduce the storage overhead in reliable distributed storage systems. Yet, such codes experience high overhead…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-06-20 Anne-Marie Kermarrec , Erwan Le Merrer , Gilles Straub , Alexandre van Kempen

Replication is a standard technique for fault tolerance in distributed systems modeled as deterministic finite state machines (DFSMs or machines). To correct f crash or f/2 Byzantine faults among n different machines, replication requires…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-03-26 Bharath Balasubramanian , Vijay K. Garg

Transformer-based language models are widely deployed for reasoning, yet their behavior under inference-time stochasticity remains underexplored. While dropout is common during training, its inference-time effects via Monte Carlo sampling…

Machine Learning · Computer Science 2026-03-19 Antônio Junior Alves Caiado , Michael Hahsler

This paper studies the problem of reconstructing a word given several of its noisy copies. This setup is motivated by several applications, among them is reconstructing strands in DNA-based storage systems. Under this paradigm, a word is…

Information Theory · Computer Science 2020-01-17 Omer Sabary , Eitan Yaakobi , Alexander Yucovich

Iterative methods are commonly used approaches to solve large, sparse linear systems, which are fundamental operations for many modern scientific simulations. When the large-scale iterative methods are running with a large number of ranks…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-05-30 Dingwen Tao , Sheng Di , Xin Liang , Zizhong Chen , Franck Cappello

Large language models (LLMs) have recently achieved significant success across various application domains, garnering substantial attention from different communities. Unfortunately, even for the best LLM, many \textit{faults} still exist…

Software Engineering · Computer Science 2024-11-06 Qiang Hu , Jin Wen , Maxime Cordy , Yuheng Huang , Wei Ma , Xiaofei Xie , Lei Ma

As large language models continue to scale up, distributed training systems have expanded beyond 10k nodes, intensifying the importance of fault tolerance. Checkpoint has emerged as the predominant fault tolerance strategy, with extensive…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-10 Weilin Cai , Le Qin , Jiayi Huang

Program errors can occur in any type of programming, and can manifest in a variety of ways, such as unexpected output, crashes, or performance issues. And program error diagnosis can often be too abstract or technical for developers to…

Software Engineering · Computer Science 2025-01-07 Zhenyu Xu , Victor S. Sheng

Corruption is notoriously widespread in data collection. Despite extensive research, the existing literature predominantly focuses on specific settings and learning scenarios, lacking a unified view of corruption modelization and…

Machine Learning · Computer Science 2026-05-19 Laura Iacovissi , Nan Lu , Robert C. Williamson

The idle computers on a local area, campus area, or even wide area network represent a significant computational resource---one that is, however, also unreliable, heterogeneous, and opportunistic. This type of resource has been used…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Adriana Iamnitchi , Ian Foster

Computing at the exascale level is expected to be affected by a significantly higher rate of faults, due to increased component counts as well as power considerations. Therefore, current day numerical algorithms need to be reexamined as to…

Numerical Analysis · Mathematics 2019-05-27 Mark Ainsworth , Christian Glusa

In a distributed storage system, recovering from multiple failures is a critical and frequent task that is crucial for maintaining the system's reliability and fault-tolerance. In this work, we focus on the problem of repairing multiple…

Information Theory · Computer Science 2018-05-09 Marwen Zorgui , Zhiying Wang