中文
相关论文

相关论文: The Life and Death of SSDs and HDDs: Similarities,…

200 篇论文

Data storage systems and their availability play a crucial role in contemporary datacenters. Despite using mechanisms such as automatic fail-over in datacenters, the role of human agents and consequently their destructive errors is…

性能 · 计算机科学 2018-06-12 Mostafa Kishani , Hossein Asadi

Dynamic random access memory failures are a threat to the reliability of data centres as they lead to data loss and system crashes. Timely predictions of memory failures allow for taking preventive measures such as server migration and…

分布式、并行与集群计算 · 计算机科学 2022-12-21 Jasmin Bogatinovski , Qiao Yu , Jorge Cardoso , Odej Kao

Existing machine learning approaches for data-driven predictive maintenance are usually black boxes that claim high predictive power yet cannot be understood by humans. This limits the ability of humans to use these models to derive…

机器学习 · 计算机科学 2021-02-15 Maxime Amram , Jack Dunn , Jeremy J. Toledano , Ying Daisy Zhuo

SSDs are emerging storage devices which unlike HDDs, do not have mechanical parts and therefore, have superior performance compared to HDDs. Due to the high cost of SSDs, entirely replacing HDDs with SSDs is not economically justified.…

性能 · 计算机科学 2018-12-12 Reza Salkhordeh , Mostafa Hadizadeh , Hossein Asadi

Neural personalized recommendation models are used across a wide variety of datacenter applications including search, social media, and entertainment. State-of-the-art models comprise large embedding tables that have billions of parameters…

硬件体系结构 · 计算机科学 2021-02-02 Mark Wilkening , Udit Gupta , Samuel Hsia , Caroline Trippel , Carole-Jean Wu , David Brooks , Gu-Yeon Wei

When will a server fail catastrophically in an industrial datacenter? Is it possible to forecast these failures so preventive actions can be taken to increase the reliability of a datacenter? To answer these questions, we have studied what…

分布式、并行与集群计算 · 计算机科学 2017-09-20 You-Luen Lee , Da-Cheng Juan , Xuan-An Tseng , Yu-Ting Chen , Shih-Chieh Chang

As the cost-per-byte of storage systems dramatically decreases, SSDs are finding their ways in emerging cloud infrastructure. Similar trend is happening for main memory subsystem, as advanced DRAM technologies with higher capacity,…

分布式、并行与集群计算 · 计算机科学 2018-08-16 Hosein Mohammadi Makrani

In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to…

分布式、并行与集群计算 · 计算机科学 2021-08-04 Bo Fang , Daoce Wang , Sian Jin , Quincey Koziol , Zhao Zhang , Qiang Guan , Suren Byna , Sriram Krishnamoorthy , Dingwen Tao

Continuous availability of HPC systems built from commodity components have become a primary concern as system size grows to thousands of processors. In this paper, we present the analysis of 8-24 months of real failure data collected from…

分布式、并行与集群计算 · 计算机科学 2013-02-21 Charng-Da Lu

NAND flash memory is ubiquitous in everyday life today because its capacity has continuously increased and cost has continuously decreased over decades. This positive growth is a result of two key trends: (1) effective process technology…

硬件体系结构 · 计算机科学 2018-01-08 Yu Cai , Saugata Ghose , Erich F. Haratsch , Yixin Luo , Onur Mutlu

Solid-state storage architectures based on NAND or emerging memory devices (SSD), are fundamentally architected and optimized for both reliability and performance. Achieving these simultaneous goals requires co-design of memory components…

硬件体系结构 · 计算机科学 2026-03-20 Jay Sarkar , Vamsi Pavan Rayaprolu , Abhijeet Bhalerao

In this paper we analyze the influence that lower layers (file system, OS, SSD) have on HDFS' ability to extract maximum performance from SSDs on the read path. We uncover and analyze three surprising performance slowdowns induced by lower…

操作系统 · 计算机科学 2019-03-25 María F. Borge , Florin Dinu , Willy Zwaenepoel

Large-scale cloud data centers have gained popularity due to their high availability, rapid elasticity, scalability, and low cost. However, current data centers continue to have high failure rates due to the lack of proper resource…

分布式、并行与集群计算 · 计算机科学 2024-12-10 Faisal Haque Bappy , Tariqul Islam , Tarannum Shaila Zaman , Raiful Hasan , Carlos Caicedo

In the current landscape of big data, the reliability and performance of storage systems are essential to the success of various applications and services. as data volumes continue to grow exponentially, the complexity and scale of the…

分布式、并行与集群计算 · 计算机科学 2025-02-06 Joshua Ludolf , Yesmin Reyna-Hernandez , Matthew Trevino

One of the most important parts of cloud computing is storage devices, and Redundant Array of Independent Disks (RAID) systems are well known and frequently used storage devices. With the increasing production of data in cloud environments,…

分布式、并行与集群计算 · 计算机科学 2021-04-06 Leila Namvari-Tazehkand , Saeid Pashazadeh

Solid state disks (SSDs) have advanced to outperform traditional hard drives significantly in both random reads and writes. However, heavy random writes trigger fre- quent garbage collection and decrease the performance of SSDs. In an SSD…

操作系统 · 计算机科学 2015-06-26 Da Zheng , Randal Burns , Alexander S. Szalay

This study investigates crash severity risk modeling strategies for work zones involving large vehicles (i.e., trucks, buses, and vans) under crash data imbalance between low-severity (LS) and high-severity (HS) crashes. We utilized crash…

机器学习 · 计算机科学 2026-02-24 Abdullah Al Mamun , Abyad Enan , Debbie A. Indah , Judith Mwakalonge , Gurcan Comert , Mashrur Chowdhury

RAID proposal advocated replacing large disks with arrays of PC disks, but as the capacity of small disks increased 100-fold in 1990s the production of large disks was discontinued. Storage dependability is increased via replication or…

分布式、并行与集群计算 · 计算机科学 2024-01-09 Alexander Thomasian

Hard Disk Drive (HDD) failures in datacenters are costly - from catastrophic data loss to a question of goodwill, stakeholders want to avoid it like the plague. An important tool in proactively monitoring against HDD failure is timely…

机器学习 · 计算机科学 2023-09-07 Rohan Mohapatra , Saptarshi Sengupta

Dataloaders, in charge of moving data from storage into GPUs while training machine learning models, might hold the key to drastically improving the performance of training jobs. Recent advances have shown promise not only by considerably…

分布式、并行与集群计算 · 计算机科学 2022-09-29 Iason Ofeidis , Diego Kiedanski , Leandros Tassiulas