中文
相关论文

相关论文: Collie: Finding Performance Anomalies in RDMA Subs…

200 篇论文

Anomaly detection is a crucial and challenging subject that has been studied within diverse research areas. In this work, we explore the task of log anomaly detection (especially computer system logs and user behavior logs) by analyzing…

机器学习 · 计算机科学 2021-01-08 Yicheng Guo , Yujin Wen , Congwei Jiang , Yixin Lian , Yi Wan

Timely and accurate detection of anomalies in power electronics is becoming increasingly critical for maintaining complex production systems. Robust and explainable strategies help decrease system downtime and preempt or mitigate…

The challenge of designing an efficient Medium Access Control (MAC) protocol and analyzing it has been an important research topic for over 30 years. This paper focuses on the performance analysis (through simulation) and modification of a…

密码学与安全 · 计算机科学 2009-07-06 Piyush Kumar Shukla , Dr. S. Silakari , Dr. Sarita Singh Bhadoria

Throughput-oriented computing via co-running multiple applications in the same machine has been widely adopted to achieve high hardware utilization and energy saving on modern supercomputers and data centers. However, efficiently co-running…

性能 · 计算机科学 2023-03-29 Hao Xu , Shuang Song , Ze Mao

GPUs are critical for compute-intensive applications, yet emerging workloads such as recommender systems, graph analytics, and data analytics often exceed GPU memory capacity. Existing solutions allow GPUs to use CPU DRAM or SSDs as…

分布式、并行与集群计算 · 计算机科学 2025-08-27 Zhuoping Yang , Jinming Zhuang , Xingzhen Chen , Alex K. Jones , Peipei Zhou

The increasing complexity and scale of telecommunication networks have led to a growing interest in automated anomaly detection systems. However, the classification of anomalies detected on network Key Performance Indicators (KPI) has…

机器学习 · 计算机科学 2023-09-01 Korantin Bordeau-Aubert , Justin Whatley , Sylvain Nadeau , Tristan Glatard , Brigitte Jaumard

The advent of RoCE (RDMA over Converged Ethernet) has led to a significant increase in the use of RDMA in datacenter networks. To achieve good performance, RoCE requires a lossless network which is in turn achieved by enabling Priority Flow…

网络与互联网体系结构 · 计算机科学 2018-06-22 Radhika Mittal , Alexander Shpiner , Aurojit Panda , Eitan Zahavi , Arvind Krishnamurthy , Sylvia Ratnasamy , Scott Shenker

Monitoring the status of large computing systems is essential to identify unexpected behavior and improve their performance and uptime. However, due to the large-scale and distributed design of such computing systems as well as a large…

分布式、并行与集群计算 · 计算机科学 2024-02-09 Tom Richard Vargis , Siavash Ghiasvand

Modern scientific workflows are data-driven and are often executed on distributed, heterogeneous, high-performance computing infrastructures. Anomalies and failures in the workflow execution cause loss of scientific productivity and…

软件工程 · 计算机科学 2021-03-24 Huy Tu , George Papadimitriou , Mariam Kiran , Cong Wang , Anirban Mandal , Ewa Deelman , Tim Menzies

Anomaly detection methods have demonstrated remarkable success across various applications. However, assessing their performance, particularly at the pixel-level, presents a complex challenge due to the severe imbalance that is most…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Mehdi Rafiei , Toby P. Breckon , Alexandros Iosifidis

Many hardware structures in today's high-performance out-of-order processors do not scale in an efficient way. To address this, different solutions have been proposed that build execution schedules in an energy-efficient manner. Issue time…

硬件体系结构 · 计算机科学 2021-09-08 Andreas Diavastos , Trevor E. Carlson

Conventional wisdom holds that an efficient interface between an OS running on a CPU and a high-bandwidth I/O device should use Direct Memory Access (DMA) to offload data transfer, descriptor rings for buffering and queuing, and interrupts…

硬件体系结构 · 计算机科学 2025-04-25 Anastasiia Ruzhanskaia , Pengcheng Xu , David Cock , Timothy Roscoe

As Deep Neural Networks (DNNs) have become an increasingly ubiquitous workload, the range of libraries and tooling available to aid in their development and deployment has grown significantly. Scalable, production quality tools are freely…

机器学习 · 计算机科学 2022-06-22 Perry Gibson , José Cano

As RISC-V architectures proliferate across embedded and high-performance domains, developers face persistent challenges in performance optimization due to fragmented tooling, immature hardware features, and platform-specific defects. This…

性能 · 计算机科学 2025-07-31 Alexander Batashev

In large-scale online services, crucial metrics, a.k.a., key performance indicators (KPIs), are monitored periodically to check their running statuses. Generally, KPIs are aggregated along multiple dimensions and derived by complex…

人工智能 · 计算机科学 2022-09-02 Shifu Yan , Caihua Shan , Wenyi Yang , Bixiong Xu , Dongsheng Li , Lili Qiu , Jie Tong , Qi Zhang

Industrial anomaly detection is an important task within computer vision with a wide range of practical use cases. The small size of anomalous regions in many real-world datasets necessitates processing the images at a high resolution. This…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Blaž Rolih , Dick Ameln , Ashwin Vaidya , Samet Akcay

The rapid growth in mobile broadband usage and increasing subscribers have made it crucial to ensure reliable network performance. As mobile networks grow more complex, especially during peak hours, manual collection of Key Performance…

网络与互联网体系结构 · 计算机科学 2024-10-08 Nooruddin Noonari , Daniel Corujo , Rui L. Aguiar , Francisco J. Ferrao

Modern manufacturers are currently undertaking the integration of novel digital technologies - such as 5G-based wireless networks, the Internet of Things (IoT), and cloud computing - to elevate their production process to a brand new level,…

网络与互联网体系结构 · 计算机科学 2021-12-08 Huanzhuo Wu , Jia He , Máté Tömösközi , Zuo Xiang , Frank H. P. Fitzek

Students interactions while solving problems in learning environments (i.e. log data) are often used to support students learning. For example, researchers use log data to develop systems that can provide students with personalized problem…

计算机与社会 · 计算机科学 2024-07-26 Alex Hicks , Yang Shi , Arun-Balajiee Lekshmi-Narayanan , Wei Yan , Samiha Marwan

As we have entered Exascale computing, the faults in high-performance systems are expected to increase considerably. To compensate for a higher failure rate, the standard checkpoint/restart technique would need to create checkpoints at a…

分布式、并行与集群计算 · 计算机科学 2023-10-26 Sarthak Joshi , Sathish Vadhiyar