中文
相关论文

相关论文: SuperBench: Improving Cloud AI Infrastructure Reli…

200 篇论文

Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system identifies, how it prioritizes them, or whether those priorities align with the review…

人工智能 · 计算机科学 2026-04-23 Ming Jin

Cloud performance diagnosis and prediction is a challenging problem due to the stochastic nature of the cloud systems. Cloud performance is affected by a large set of factors including (but not limited to) virtual machine types, regions,…

分布式、并行与集群计算 · 计算机科学 2016-12-19 Karan Mitra , Saguna Saguna , Christer Åhlund , Rajiv Ranjan

The majority of cloud providers offers users the possibility to deploy Trusted Execution Environments (TEEs) to protect their data and processes from high privileged adversaries. This offer is intended to address concerns of users when…

密码学与安全 · 计算机科学 2022-05-09 Mathias Morbitzer , Benedikt Kopf , Philipp Zieris

Today, Neural Networks are the basis of breakthroughs in virtually every technical domain. Their application to accelerators has recently resulted in better performance and efficiency in these systems. At the same time, the increasing…

机器学习 · 计算机科学 2021-12-07 Jashanpreet Singh Sraw , Deepak M C

The opacity of AI models necessitates both validation and evaluation before their integration into services. To investigate these models, explainable AI (XAI) employs methods that elucidate the relationship between input features and output…

密码学与安全 · 计算机科学 2024-10-02 Zerui Wang , Yan Liu

Trends such as cloud computing raise issues regarding stable and uniform quality assurance and validation of software requirements. Current QA frameworks are poorly defined, often not automated, and lack the flexibility needed for…

软件工程 · 计算机科学 2025-02-20 Mohammed Alharbi , RJ Qureshi

How well do AI systems perform in algorithm engineering for hard optimization problems in domains such as package-delivery routing, crew scheduling, factory production planning, and power-grid balancing? We introduce ALE-Bench, a new…

人工智能 · 计算机科学 2025-10-07 Yuki Imajuku , Kohki Horie , Yoichi Iwata , Kensho Aoki , Naohiro Takahashi , Takuya Akiba

High-Performance Computing (HPC) centers and cloud providers support an increasingly diverse set of applications on heterogenous hardware. As Artificial Intelligence (AI) and Machine Learning (ML) workloads have become an increasingly…

Enhancing the reliability of AI based fault diagnosis in inverter dominated microgrids requires diverse and statistically balanced datasets. However, the scarcity and imbalance of high fidelity fault data, especially for rare inverter…

系统与控制 · 电气工程与系统科学 2025-11-11 Swetha Rani Kasimalla , Kuchan Park , Junho Hong , Young-Jin Kim

The deployment of deep learning inference in production environments continues to grow, where throughput, latency, and hardware efficiency are critical. Although specialized accelerators are increasingly adopted, many inference workloads…

性能 · 计算机科学 2026-02-20 Kathiravan Palaniappan

As artificial intelligence (AI) continues to permeate various domains, concerns surrounding trust and transparency in AI-driven inference and training processes have emerged, particularly with respect to potential biases and traceability…

分布式、并行与集群计算 · 计算机科学 2023-05-09 Sanghyeon Park , Junmo Lee , Soo-Mook Moon

Cloud computing systems fail in complex and unexpected ways due to unexpected combinations of events and interactions between hardware and software components. Fault injection is an effective means to bring out these failures in a…

软件工程 · 计算机科学 2020-10-02 Domenico Cotroneo , Luigi De Simone , Pietro Liguori , Roberto Natella

AI systems can fail to learn important behaviors, leading to real-world issues like safety concerns and biases. Discovering these systematic failures often requires significant developer attention, from hypothesizing potential edge cases to…

人机交互 · 计算机科学 2021-10-28 Ángel Alexander Cabrera , Abraham J. Druck , Jason I. Hong , Adam Perer

Performance unpredictability is a major roadblock towards cloud adoption, and has performance, cost, and revenue ramifications. Predictable performance is even more critical as cloud services transition from monolithic designs to…

分布式、并行与集群计算 · 计算机科学 2019-05-06 Yu Gan , Yanqi Zhang , Kelvin Hu , Dailun Cheng , Yuan He , Meghna Pancholi , Christina Delimitrou

We present a blockchain based system that allows data owners, cloud vendors, and AI developers to collaboratively train machine learning models in a trustless AI marketplace. Data is a highly valued digital asset and central to deriving…

分布式、并行与集群计算 · 计算机科学 2020-02-04 Nishant Baranwal Somy , Kalapriya Kannan , Vijay Arya , Sandeep Hans , Abhishek Singh , Pranay Lohia , Sameep Mehta

Adaptive medical AI models often face performance drops in dynamic clinical environments due to data drift. We propose an autonomous continuous monitoring and data integration framework that maintains robust performance over time. Focusing…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Mohammad Daouk , Jan Ulrich Becker , Neeraja Kambham , Anthony Chang , Chandra Mohan , Hien Van Nguyen

We conduct an empirical study of machine learning functionalities provided by major cloud service providers, which we call machine learning clouds. Machine learning clouds hold the promise of hiding all the sophistication of running…

分布式、并行与集群计算 · 计算机科学 2017-10-17 Yu Liu , Hantian Zhang , Luyuan Zeng , Wentao Wu , Ce Zhang

Misconfiguration, excessive privilege, and fragmented controls remain major causes of cloud-infrastructure incidents. This paper proposes an open-source framework that contributes a cross-platform identity-resource graph for Kubernetes and…

密码学与安全 · 计算机科学 2026-04-29 Wanru Shao

Affordances and permissions are promising and timely safety levers for mitigating Loss of Control (LoC) threats in high-stakes deployment contexts, such as national security. Deployers in defense and intelligence could rely on several…

计算机与社会 · 计算机科学 2026-05-21 Matteo Pistillo , Samantha Faraone , Joshua Herman

Context: Blockchain and AI are increasingly explored to enhance trustworthiness in software engineering (SE), particularly in supporting software evolution tasks. Method: We conducted a systematic literature review (SLR) using a predefined…

软件工程 · 计算机科学 2026-02-03 Mohammad Naserameri , Juergen Rilling