中文
相关论文

相关论文: TSGuard: Automated User-Centric Incident Diagnosis…

200 篇论文

In a hierarchically-structured cloud/edge/device computing environment, workload allocation can greatly affect the overall system performance. This paper deals with AI-oriented medical workload generated in emergency rooms (ER) or intensive…

分布式、并行与集群计算 · 计算机科学 2020-02-11 Tianshu Hao , Jianfeng Zhan , Kai Hwang , Wanling Gao , Xu Wen

In the construction industry, safety assessment is vital to ensure both the reliability of assets and the safety of workers. Scaffolding, a key structural support asset requires regular inspection to detect and identify alterations from the…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Sameer Prabhu , Amit Patwardhan , Ramin Karim

Fault diagnosis has attracted extensive attention for its importance in the exceedingly fault management framework for cloud virtualization, despite the fact that fault diagnosis becomes more difficult due to the increasing scalability and…

软件工程 · 计算机科学 2015-07-30 Ameen Alkasem , Hongwei Liu , Zuo Decheng , Yao Zhao

The proliferation of AI technology gives rise to a variety of security threats, which significantly compromise the confidentiality and integrity of AI models and applications. Existing software-based solutions mainly target one specific…

密码学与安全 · 计算机科学 2023-11-29 Xiaobei Yan , Han Qiu , Tianwei Zhang

Public EV charging infrastructure suffers from significant failure rates -- with field studies reporting up to 27.5% of DC fast chargers non-functional -- and multi-day mean time to resolution, imposing billions in annual economic burden.…

分布式、并行与集群计算 · 计算机科学 2026-03-11 Mohammed Cherifi

Cloud application services are distributed in nature and have components across the stack working together to deliver the experience to end users. The wide adoption of microservice architecture exacerbates failure management due to…

性能 · 计算机科学 2025-09-09 Dhanya R Mathews , Mudit Verma , Pooja Aggarwal , J. Lakshmi

Modern machine learning workloads such as large language model training, fine-tuning jobs are highly distributed and span across hundreds of systems with multiple GPUs. Job completion time for these workloads is the artifact of the…

网络与互联网体系结构 · 计算机科学 2025-07-08 Jit Gupta , Tarun Banka , Rahul Gupta , Mithun Dharmaraj , Jasleen Kaur

Incidents in microservice environments can be costly and challenging to recover from due to their complexity and distributed nature. Recent advancements in artificial intelligence (AI) offer promising solutions for improving incident…

软件工程 · 计算机科学 2024-10-08 Dahlia Ziqi Zhou , Marios Fokaefs

As artificial intelligence (AI) systems become increasingly deployed across the world, they are also increasingly implicated in AI incidents - harm events to individuals and society. As a result, industry, civil society, and governments…

计算机与社会 · 计算机科学 2024-09-26 Kevin Paeth , Daniel Atherton , Nikiforos Pittaras , Heather Frase , Sean McGregor

The ever-increasing demand for generative artificial intelligence (GenAI) has motivated cloud-based GenAI services such as Azure OpenAI Service and Amazon Bedrock. Like any large-scale cloud service, failures are inevitable in cloud-based…

分布式、并行与集群计算 · 计算机科学 2025-08-15 Haoran Yan , Yinfang Chen , Minghua Ma , Ming Wen , Shan Lu , Shenglin Zhang , Tianyin Xu , Rujia Wang , Chetan Bansal , Saravan Rajmohan , Qingwei Lin , Chaoyun Zhang , Dongmei Zhang

Real-time detection and mitigation of technical anomalies are critical for large-scale cloud-native services, where even minutes of downtime can result in massive financial losses and diminished user trust. While customer incidents serve as…

计算与语言 · 计算机科学 2026-05-25 Jun Wang , Ziyin Zhang , Rui Wang , Hang Yu , Peng Di , Rui Wang

As the modern microservice architecture for cloud applications grows in popularity, cloud services are becoming increasingly complex and more vulnerable to misconfiguration and software bugs. Traditional approaches rely on expert input to…

软件工程 · 计算机科学 2026-05-21 Rohan Kumar , Jason Li , Zongshun Zhang , Syed Mohammad Qasim , Gianluca Stringhini , Ayse K. Coskun

Modern software engineers operate across 5-10 disconnected tools daily: GitHub, GitLab, Jira, Slack, calendar applications, CI dashboards, AI coding assistants, and container platforms. This fragmentation creates cognitive overhead that…

软件工程 · 计算机科学 2026-04-21 Happy Bhati

The reliability of cloud platforms is of significant relevance because society increasingly relies on complex software systems running on the cloud. To improve it, cloud providers are automating various maintenance tasks, with failure…

软件工程 · 计算机科学 2022-04-07 Jasmin Bogatinovski , Sasho Nedelkoski , Li Wu , Jorge Cardoso , Odej Kao

End-point monitoring solutions are widely deployed in today's enterprise environments to support advanced attack detection and investigation. These monitors continuously record system-level activities as audit logs and provide deep…

密码学与安全 · 计算机科学 2026-02-16 Hao Zhang , Shuo Shao , Song Li , Zhenyu Zhong , Yan Liu , Zhan Qin

Operating Elasticsearch clusters at scale demands continuous human expertise spanning the full lifecycle -- from initial deployment through performance tuning, monitoring, failure prediction, and incident recovery. We present the ES…

分布式、并行与集群计算 · 计算机科学 2026-04-07 Muhamed Ramees Cheriya Mukkolakkal

Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must…

机器学习 · 计算机科学 2025-10-27 Stefan Kraft , Andreas Theissler , Vera Wienhausen-Wilke , Gjergji Kasneci , Hendrik Lensch

The AI Incident Database was inspired by aviation safety databases, which enable collective learning from failures to prevent future incidents. The database documents hundreds of AI failures, collected from the news and media. However,…

计算机与社会 · 计算机科学 2025-05-08 Isabel Richards , Claire Benn , Miri Zilka

The rapid evolution of Artificial Intelligence (AI) and Machine Learning (ML) has significantly heightened computational demands, particularly for inference-serving workloads. While traditional cloud-based deployments offer scalability,…

分布式、并行与集群计算 · 计算机科学 2025-09-17 Foteini Stathopoulou , Aggelos Ferikoglou , Manolis Katsaragakis , Dimosthenis Masouros , Sotirios Xydis , Dimitrios Soudris

Existing processes and methods for incident handling are geared towards infrastructures and operational models that will be increasingly outdated by cloud computing. Research has shown that to adapt incident handling to cloud computing…

密码学与安全 · 计算机科学 2016-06-28 Duane Wilson , Jeff Avery