中文
相关论文

相关论文: High Significant Fault Detection in Azure Core Wor…

200 篇论文

Detecting point anomalies in bank account balances is essential for financial institutions, as it enables the identification of potential fraud, operational issues, or other irregularities. Robust statistics is useful for flagging outliers…

机器学习 · 计算机科学 2025-12-02 Federico Maddanu , Tommaso Proietti , Riccardo Crupi

Microsoft Azure is dedicated to guarantee high quality of service to its customers, in particular, during periods of high customer activity, while controlling cost. We employ a Data Science (DS) driven solution to predict user load and…

Observations in data which are significantly different from its neighbouring points but cannot be classified as noise are known as anomalies or outliers. These anomalies are a cause of concern and a timely warning about their presence could…

应用统计 · 统计学 2020-06-09 Krishnam Kapoor

The early detection of anomalous events in time series data is essential in many domains of application. In this paper we deal with critical health events, which represent a significant cause of mortality in intensive care units of…

机器学习 · 统计学 2020-10-23 Vitor Cerqueira , Luis Torgo , Carlos Soares

The complexity and diversity of big data and AI workloads make understanding them difficult and challenging. This paper proposes a new approach to characterizing big data and AI workloads. We consider each big data and AI workload as a…

分布式、并行与集群计算 · 计算机科学 2018-05-07 Wanling Gao , Jianfeng Zhan , Lei Wang , Chunjie Luo , Daoyi Zheng , Fei Tang , Biwei Xie , Chen Zheng , Qiang Yang

Power-generating assets (e.g., jet engines, gas turbines) are often instrumented with tens to hundreds of sensors for monitoring physical and performance degradation. Anomaly detection algorithms highlight deviations from predetermined…

分布式、并行与集群计算 · 计算机科学 2017-01-27 Paras Jain , Chirag Tailor , Sam Ford , Liexiao Ding , Michael Phillips , Fang Liu , Nagi Gebraeel , Duen Horng Chau

Within today's large-scale systems, one anomaly can impact millions of users. Detecting such events in real-time is essential to maintain the quality of services. It allows the monitoring team to prevent or diminish the impact of a failure.…

人工智能 · 计算机科学 2023-04-25 Arthur Vervaet

Reliability is a cumbersome problem in High Performance Computing Systems and Data Centers evolution. During operation, several types of fault conditions or anomalies can arise, ranging from malfunctioning hardware to improper…

分布式、并行与集群计算 · 计算机科学 2020-07-30 Andrea Borghesi , Antonio Libri , Luca Benini , Andrea Bartolini

Detecting anomalies in time series data is important in a variety of fields, including system monitoring, healthcare, and cybersecurity. While the abundance of available methods makes it difficult to choose the most appropriate method for a…

机器学习 · 计算机科学 2023-02-03 Ferdinand Rewicki , Joachim Denzler , Julia Niebling

As contemporary software-intensive systems reach increasingly large scale, it is imperative that failure detection schemes be developed to help prevent costly system downtimes. A promising direction towards the construction of such schemes…

应用统计 · 统计学 2016-09-27 Alexey Artemov , Evgeny Burnaev

Azure Cosmos DB is a cloud-native distributed database, operating at a massive scale, powering Microsoft Cloud. Think 10s of millions of database partitions (replica-sets), 100+ PBs of data under management, 20M+ vCores. Failovers are an…

Time series anomaly detection is crucial for industrial monitoring services that handle a large volume of data, aiming to ensure reliability and optimize system performance. Existing methods often require extensive labeled resources and…

机器学习 · 计算机科学 2023-07-21 Manqing Dong , Zhanxiang Zhao , Yitong Geng , Wentao Li , Wei Wang , Huai Jiang

Beyond implementation correctness of a distributed system, it is equally important to understand exactly what users should expect to see from that system. Even if the system itself works as designed, insufficient understanding of its…

软件工程 · 计算机科学 2022-10-26 A. Finn Hackett , Joshua Rowe , Markus Alexander Kuppe

Anomaly detection in distributed systems such as High-Performance Computing (HPC) clusters is vital for early fault detection, performance optimisation, security monitoring, reliability in general but also operational insights. Deep Neural…

机器学习 · 计算机科学 2024-05-14 Franz Kevin Stehle , Wainer Vandelli , Giuseppe Avolio , Felix Zahn , Holger Fröning

Anomaly detection (AD) is a fundamental task for time-series analytics with important implications for the downstream performance of many applications. In contrast to other domains where AD mainly focuses on point-based anomalies (i.e.,…

Low latency and high availability of an app or a web service are key, amongst other factors, to the overall user experience (which in turn directly impacts the bottomline). Exogenic and/or endogenic factors often give rise to breakouts in…

统计方法学 · 统计学 2014-12-01 Nicholas A. James , Arun Kejariwal , David S. Matteson

Anomaly detection is the task of identifying examples that do not behave as expected. Because anomalies are rare and unexpected events, collecting real anomalous examples is often challenging in several applications. In addition, learning…

机器学习 · 计算机科学 2024-05-24 Lorenzo Perini , Maja Rudolph , Sabrina Schmedding , Chen Qiu

Workloads in modern cloud data centers are becoming increasingly complex. The number of workloads running in cloud data centers has been growing exponentially for the last few years, and cloud service providers (CSP) have been supporting…

分布式、并行与集群计算 · 计算机科学 2022-11-30 Mohammad Hossain , Derssie Mebratu , Niranjan Hasabnis , Jun Jin , Gaurav Chaudhary , Noah Shen

Network troubleshooting is still a heavily human-intensive process. To reduce the time spent by human operators in the diagnosis process, we present a system based on (i) unsupervised learning methods for detecting anomalies in the time…

网络与互联网体系结构 · 计算机科学 2021-08-27 Jose M. Navarro , Alexis Huet , Dario Rossi

Several techniques for multivariate time series anomaly detection have been proposed recently, but a systematic comparison on a common set of datasets and metrics is lacking. This paper presents a systematic and comprehensive evaluation of…

机器学习 · 计算机科学 2021-09-24 Astha Garg , Wenyu Zhang , Jules Samaran , Savitha Ramasamy , Chuan-Sheng Foo