中文
相关论文

相关论文: DRAM Failure Prediction in AIOps: Empirical Evalua…

200 篇论文

Modern datacenters assemble a very large number of disk drives under a single roof. Even if economic and technical factors where to make individual drives more reliable (which is not at all clear, given the commoditization of the…

分布式、并行与集群计算 · 计算机科学 2017-07-10 Jayanta Basak , Randy H. Katz

As software systems grow increasingly intricate, Artificial Intelligence for IT Operations (AIOps) methods have been widely used in software system failure management to ensure the high availability and reliability of large-scale…

软件工程 · 计算机科学 2024-06-25 Lingzhe Zhang , Tong Jia , Mengxi Jia , Yifan Wu , Aiwei Liu , Yong Yang , Zhonghai Wu , Xuming Hu , Philip S. Yu , Ying Li

The use of credit cards has recently increased, creating an essential need for credit card assessment methods to minimize potential risks. This study investigates the utilization of machine learning (ML) models for credit card default…

机器学习 · 计算机科学 2023-10-17 Anas Arram , Masri Ayob , Musatafa Abbas Abbood Albadr , Alaa Sulaiman , Dheeb Albashish

Advancement in Processor technology has made it easy to handle data-intensive workloads, but limiting main memory advances has created performance bottlenecks. In DRAM, there have been improvements in DRAM access latency as well as…

硬件体系结构 · 计算机科学 2021-05-24 Saurabh Jaiswal , Shailendra Kumar Gupta , Soumya Soubhagya Dandapat

With the rapid development of cloud computing and big data technologies, storage systems have become a fundamental building block of datacenters, incorporating hardware innovations such as flash solid state drives and non-volatile memories,…

数据库 · 计算机科学 2023-08-01 Chenyuan Wu

A classification technique incorporating a novel feature derivation method is proposed for predicting failure of a system or device with multivariate time series sensor data. We treat the multivariate time series sensor data as images for…

机器学习 · 计算机科学 2021-09-22 Lanfa Frank Wang , Danjue Li

Artificial intelligence (AI) and Machine Learning (ML) are becoming pervasive in today's applications, such as autonomous vehicles, healthcare, aerospace, cybersecurity, and many critical applications. Ensuring the reliability and…

硬件体系结构 · 计算机科学 2021-03-31 Shamik Kundu , Kanad Basu , Mehdi Sadi , Twisha Titirsha , Shihao Song , Anup Das , Ujjwal Guin

Functional verification and debugging are critical bottlenecks in modern System-on-Chip (SoC) design, with manual detection of Advanced Peripheral Bus (APB) transaction errors in large Value Change Dump (VCD) files being inefficient and…

软件工程 · 计算机科学 2025-09-05 Cheng-Yang Tsai , Tzu-Wei Huang , Jen-Wei Shih , I-Hsiang Wang , Yu-Cheng Lin , Rung-Bin Lin

Artificial Intelligence for IT Operations (AIOps) has been adopted in organizations in various tasks, including interpreting models to identify indicators of service failures. To avoid misleading practitioners, AIOps model interpretations…

机器学习 · 计算机科学 2022-02-07 Yingzhe Lyu , Gopi Krishnan Rajbahadur , Dayi Lin , Boyuan Chen , Zhen Ming , Jiang

Modern machine learning training is increasingly bottlenecked by data I/O rather than compute. GPUs often sit idle at below 50% utilization waiting for data. This paper presents a machine learning approach to predict I/O performance and…

性能 · 计算机科学 2025-12-22 Karthik Prabhakar , Durgamadhab Mishra

Cloud providers are concerned that Rowhammer poses a potentially critical threat to their servers, yet today they lack a systematic way to test whether the DRAM used in their servers is vulnerable to Rowhammer attacks. This paper presents…

密码学与安全 · 计算机科学 2020-03-11 Lucian Cojocar , Jeremie Kim , Minesh Patel , Lillian Tsai , Stefan Saroiu , Alec Wolman , Onur Mutlu

When will a server fail catastrophically in an industrial datacenter? Is it possible to forecast these failures so preventive actions can be taken to increase the reliability of a datacenter? To answer these questions, we have studied what…

分布式、并行与集群计算 · 计算机科学 2017-09-20 You-Luen Lee , Da-Cheng Juan , Xuan-An Tseng , Yu-Ting Chen , Shih-Chieh Chang

Motivated by cloud computing applications, we study the problem of how to optimally deploy new hardware subject to both power and robustness constraints. To model the situation observed in large-scale data centers, we introduce the Online…

数据结构与算法 · 计算机科学 2022-09-05 Konstantina Mellou , Marco Molinaro , Rudy Zhou

The precise estimation of resource usage is a complex and challenging issue due to the high variability and dimensionality of heterogeneous service types and dynamic workloads. Over the last few years, the prediction of resource usage and…

分布式、并行与集群计算 · 计算机科学 2023-02-07 Deepika Saxena , Jitendra Kumar , Ashutosh Kumar Singh , Stefan Schmid

Software defect prediction is a critical aspect of software quality assurance, as it enables early identification and mitigation of defects, thereby reducing the cost and impact of software failures. Over the past few years, quantum…

软件工程 · 计算机科学 2024-12-11 Md Nadim , Mohammad Hassan , Ashis Kumar Mandal , Chanchal K. Roy

Existing machine learning approaches for data-driven predictive maintenance are usually black boxes that claim high predictive power yet cannot be understood by humans. This limits the ability of humans to use these models to derive…

机器学习 · 计算机科学 2021-02-15 Maxime Amram , Jack Dunn , Jeremy J. Toledano , Ying Daisy Zhuo

To provide proactive fault tolerance for modern cloud data centers, extensive studies have proposed machine learning (ML) approaches to predict imminent disk failures for early remedy and evaluated their approaches directly on public…

机器学习 · 计算机科学 2019-12-23 Shujie Han , Jun Wu , Erci Xu , Cheng He , Patrick P. C. Lee , Yi Qiang , Qixing Zheng , Tao Huang , Zixi Huang , Rui Li

Artificial Intelligence has gained a lot of attention recently, it has been utilized in several fields ranging from daily life activities, such as responding to emails and scheduling appointments, to manufacturing and automating work…

软件工程 · 计算机科学 2026-02-02 Mohammed O. Alannsary

In this paper we examine historical failures of artificial intelligence (AI) and propose a classification scheme for categorizing future failures. By doing so we hope that (a) the responses to future failures can be improved through…

计算机与社会 · 计算机科学 2019-07-19 Peter J. Scott , Roman V. Yampolskiy

Operationalizing machine learning based security detections is extremely challenging, especially in a continuously evolving cloud environment. Conventional anomaly detection does not produce satisfactory results for analysts that are…

密码学与安全 · 计算机科学 2017-09-22 Ram Shankar Siva Kumar , Andrew Wicker , Matt Swann