English
Related papers

Related papers: High Significant Fault Detection in Azure Core Wor…

200 papers

Organizations rely heavily on time series metrics to measure and model key aspects of operational and business performance. The ability to reliably detect issues with these metrics is imperative to identifying early indicators of major…

Machine Learning · Computer Science 2020-11-11 Sayan Chakraborty , Smit Shah , Kiumars Soltani , Anna Swigart , Luyao Yang , Kyle Buckingham

Large companies need to monitor various metrics (for example, Page Views and Revenue) of their applications and services in real time. At Microsoft, we develop a time-series anomaly detection service which helps customers to monitor the…

Machine Learning · Computer Science 2019-06-11 Hansheng Ren , Bixiong Xu , Yujing Wang , Chao Yi , Congrui Huang , Xiaoyu Kou , Tony Xing , Mao Yang , Jie Tong , Qi Zhang

AI workloads incur frequent failures and incidents from the underlying infrastructure. The current incident management workflow follows a provider-centric paradigm, where users report incidents to the infrastructure provider who then…

Software Engineering · Computer Science 2026-05-08 Yitao Yang , Yangtao Deng , Yifan Xiong , Baochun Li , Hong Xu , Peng Cheng

Time-series anomaly detection, which detects errors and failures in a workflow, is one of the most important topics in real-world applications. The purpose of time-series anomaly detection is to reduce potential damages or losses. However,…

Machine Learning · Computer Science 2025-04-17 Jinsung Jeon , Jaehyeon Park , Sewon Park , Jeongwhan Choi , Minjung Kim , Noseong Park

Multivariate anomaly detection can be used to identify outages within large volumes of telemetry data for computing systems. However, developing an efficient anomaly detector that can provide users with relevant information is a challenging…

Machine Learning · Computer Science 2022-02-15 Bruno Wassermann , David Ohana , Ronen Schaffer , Robert Shahla , Elliot K. Kolodner , Eran Raichstein , Michal Malka

Cloud platforms, under the hood, consist of a complex inter-connected stack of hardware and software components. Each of these components can fail which may lead to an outage. Our goal is to improve the quality of Cloud services through…

Software Engineering · Computer Science 2021-02-12 Mohammad Saiful Islam , Andriy Miranskyy

Performance and high availability have become increasingly important drivers, amongst other drivers, for user retention in the context of web services such as social networks, and web search. Exogenic and/or endogenic factors often give…

Machine Learning · Computer Science 2017-04-26 Jordan Hochenbaum , Owen S. Vallis , Arun Kejariwal

Cloud computing is ubiquitous: more and more companies are moving the workloads into the Cloud. However, this rise in popularity challenges Cloud service providers, as they need to monitor the quality of their ever-growing offerings…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-08-04 Mohammad Saiful Islam , William Pourmajidi , Lei Zhang , John Steinbacher , Tony Erwin , Andriy Miranskyy

Anomaly detection is a branch of data analysis and machine learning which aims at identifying observations that exhibit abnormal behaviour. Be it measurement errors, disease development, severe weather, production quality default(s) (items)…

Machine Learning · Statistics 2024-07-11 Pavlo Mozharovskyi , Romain Valla

Anomaly detection to recognize unusual events in large scale systems in a time sensitive manner is critical in many industries, eg. bank fraud, enterprise systems, medical alerts, etc. Large-scale systems often grow in size and complexity…

Machine Learning · Computer Science 2022-10-31 Srishti Mishra , Tvarita Jain , Dinkar Sitaram

Detecting and resolving performance anomalies in Cloud services is crucial for maintaining desired performance objectives. Scaling actions triggered by an anomaly detector help achieve target latency at the cost of extra resource…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-24 Gabriel Job Antunes Grabher , Fumio Machida , Thomas Ropars

On-line detection of anomalies in time series is a key technique used in various event-sensitive scenarios such as robotic system monitoring, smart sensor networks and data center security. However, the increasing diversity of data sources…

Machine Learning · Computer Science 2021-04-26 Wentai Wu , Ligang He , Weiwei Lin , Yi Su , Yuhua Cui , Carsten Maple , Stephen Jarvis

Deviations from expected behavior during runtime, known as anomalies, have become more common due to the systems' complexity, especially for microservices. Consequently, analyzing runtime monitoring data, such as logs, traces for…

Software Engineering · Computer Science 2024-08-16 Monika Steidl , Benedikt Dornauer , Michael Felderer , Rudolf Ramler , Mircea-Cristian Racasan , Marko Gattringer

This study proposes an anomaly detection method based on the Transformer architecture with integrated multiscale feature perception, aiming to address the limitations of temporal modeling and scale-aware feature representation in cloud…

Machine Learning · Computer Science 2025-08-26 Lian Lian , Yilin Li , Song Han , Renzi Meng , Sibo Wang , Ming Wang

This paper introduces a scalable Anomaly Detection Service with a generalizable API tailored for industrial time-series data, designed to assist Site Reliability Engineers (SREs) in managing cloud infrastructure. The service enables…

Machine Learning · Computer Science 2025-01-29 Nimesh Jha , Shuxin Lin , Srideepika Jayaraman , Kyle Frohling , Christodoulos Constantinides , Dhaval Patel

Data centers play a key role in today's Internet. Cloud applications are mainly hosted on multi-tenant warehouse-scale data centers. Anomalies pose a serious threat to data centers' operations. If not controlled properly, a simple anomaly…

Networking and Internet Architecture · Computer Science 2019-06-18 Ashkan Aghdai , Kang Xi , H. Jonathan Chao

Time series anomaly detection is usually formulated as finding outlier data points relative to some usual data, which is also an important problem in industry and academia. To ensure systems working stably, internet companies, banks and…

Machine Learning · Computer Science 2018-12-24 Zhang Rong , Dong Shandong , Nie Xin , Xiao Shiguang

Cloud systems are susceptible to performance issues, which may cause service-level agreement violations and financial losses. In current practice, crucial metrics are monitored periodically to provide insight into the operational status of…

Machine Learning · Computer Science 2024-11-08 Wenwei Gu , Jinyang Liu , Zhuangbin Chen , Jianping Zhang , Yuxin Su , Jiazhen Gu , Cong Feng , Zengyin Yang , Yongqiang Yang , Michael Lyu

Anomaly detection is concerned with identifying data patterns that deviate remarkably from the expected behaviour. This is an important research problem, due to its broad set of application domains, from data analysis to e-health,…

Machine Learning · Computer Science 2021-08-23 L. Erhan , M. Ndubuaku , M. Di Mauro , W. Song , M. Chen , G. Fortino , O. Bagdasar , A. Liotta

Online unsupervised detection of anomalies is crucial to guarantee the correct operation of cyber-physical systems and the safety of humans interacting with them. State-of-the-art approaches based on deep learning via neural networks…

Machine Learning · Computer Science 2024-07-30 Daniele Meli
‹ Prev 1 2 3 10 Next ›