中文
相关论文

相关论文: Empirical Measurements of Disk Failure Rates and E…

200 篇论文

To face future reliability challenges, it is necessary to quantify the risk of error in any part of a computing system. To this goal, the Architectural Vulnerability Factor (AVF) has long been used for chips. However, this metric is used…

硬件体系结构 · 计算机科学 2023-08-02 Luc Jaulmes , Miquel Moretó , Mateo Valero , Marc Casas

We have been investigating the use of low-cost, commodity components for multi-terabyte SQL Server databases. Dubbed storage bricks, these servers are white box PCs containing the largest ATA drives, value-priced AMD or Intel processors,…

数据库 · 计算机科学 2007-05-23 Tom Barclay , Wyman Chong , Jim Gray

We found that a reliability model commonly used to estimate Mean-Time-To-Data-Loss (MTTDL), while suitable for modeling RAID 0 and RAID 5, fails to accurately model systems having a fault-tolerance greater than 1. Therefore, to model the…

分布式、并行与集群计算 · 计算机科学 2013-10-18 Jason Resch , Ilya Volvovski

In large-scale datacenters, memory failure is a common cause of server crashes, with Uncorrectable Errors (UEs) being a major indicator of Dual Inline Memory Module (DIMM) defects. Existing approaches primarily focus on predicting UEs using…

硬件体系结构 · 计算机科学 2023-12-19 Qiao Yu , Wengui Zhang , Jorge Cardoso , Odej Kao

The aggressive scaling of technology may have helped to meet the growing demand for higher memory capacity and density, but has also made DRAM cells more prone to errors. Such a reality triggered a lot of interest in modeling DRAM behavior…

分布式、并行与集群计算 · 计算机科学 2020-03-30 Lev Mukhanov , Konstantinos Tovletoglou , Hans Vandierendonck , Dimitrios S. Nikolopoulos , Georgios Karakonstantis

The tolerable erasure error rate for scalable quantum computation is shown to be at least 0.292, given standard scalability assumptions. This bound is obtained by implementing computations with generic stabilizer code teleportation steps…

量子物理 · 物理学 2007-05-23 E. Knill

In recent years, high availability and reliability of Data Storage Systems (DSS) have been significantly threatened by soft errors occurring in storage controllers. Due to their specific functionality and hardware-software stack, error…

性能 · 计算机科学 2021-12-24 Mostafa Kishani , Mehdi Tahoori , Hossein Asadi

One of the primary objectives of a distributed storage system is to reliably store a large amount $dsize$ of source data for a long duration using a large number $N$ of unreliable storage nodes, each with capacity $nsize$. The storage…

信息论 · 计算机科学 2020-02-20 Michael Luby

Modern DRAM chips are subject to read disturbance errors. State-of-the-art read disturbance mitigations rely on accurate and exhaustive characterization of the read disturbance threshold (RDT) (e.g., the number of aggressor row activations…

The widespread prevalence of data breaches amplifies the importance of auditing storage systems. In this work, we initiate the study of auditable storage emulations, which provide the capability for an auditor to report the previously…

分布式、并行与集群计算 · 计算机科学 2020-05-19 Vinicius V. Cogo , Alysson Bessani

In order to achieve fault tolerance, highly reliable system often require the ability to detect errors as soon as they occur and prevent the speared of erroneous information throughout the system. Thus, the need for codes capable of…

信息论 · 计算机科学 2010-02-08 Muzhir Al-Ani , Qeethara Al-Shayea

State-of-the-art DRAM read disturbance mitigations rely on the read disturbance threshold (RDT) (e.g., the number of aggressor row activations needed to induce the first read disturbance bitflip) to securely and performance- and…

硬件体系结构 · 计算机科学 2026-03-16 Ataberk Olgun , F. Nisa Bostanci , Ismail Emir Yuksel , Haocong Luo , Minesh Patel , A. Giray Yaglikci , Onur Mutlu

Too many defective compute chips are escaping existing manufacturing tests -- at least an order of magnitude more than industrial targets across all compute chip types in data centers. Silent data corruptions (SDCs) caused by test escapes,…

Magnetic data storage is pervasive in the preservation of digital information and the rapid pace of computer development requires ever more capacity. Increasing the storage density for magnetic hard disk drives requires a reduced bit size,…

材料科学 · 物理学 2013-10-24 R. F. L. Evans , R. W. Chantrell , U. Nowak , A. Lyberatos , H-J. Richter

Fault tolerance is a key factor of industrial computing systems design. But in practical terms, these systems, like every commercial product, are under great financial constraints and they have to remain in operational state as long as…

系统与控制 · 计算机科学 2015-03-31 Andrey A. Shchurov

We study the problem of measuring errors in non-trace-preserving quantum operations, with a focus on their impact on quantum computing. We propose an error metric that efficiently provides an upper bound on the trace distance between the…

量子物理 · 物理学 2023-10-10 Yu Shi , Edo Waks

Solid-State Drives (SSDs) are recently employed in enterprise servers and high-end storage systems in order to enhance performance of storage subsystem. Although employing high speed SSDs in the storage subsystems can significantly improve…

其他计算机科学 · 计算机科学 2018-05-02 Saba Ahmadian , Farhad Taheri , Mehrshad Lotfi , Maryam Karimi , Hossein Asad

Robust qubit memory is essential for quantum computing, both for near-term devices operating without error correction, and for the long-term goal of a fault-tolerant processor. We directly measure the memory error $\epsilon_m$ for a…

With the rapid advancements of deep learning in recent years, hardware accelerators are continuously deployed in more and more safety-critical applications such as autonomous driving and robotics. While the accelerators are usually…

硬件体系结构 · 计算机科学 2023-08-31 Zuodong Zhang , Renjie Wei , Meng Li , Yibo Lin , Runsheng Wang , Ru Huang

High-performance and safety-critical system architects must accurately evaluate the application-level silent data corruption (SDC) rates of processors to soft errors. Such an evaluation requires error propagation all the way from particle…

‹ 上一页 1 2 3 10 下一页 ›