English
Related papers

Related papers: An Extensible Software Transport Layer for GPU Net…

200 papers

Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of megabytes at 1024-application fanout -…

Artificial Intelligence · Computer Science 2026-05-28 Bojie Li

Geometric Representation Learning (GRL) aims to approximate the non-Euclidean topology of high-dimensional data through discrete graph structures, grounded in the manifold hypothesis. However, traditional static graph construction methods…

Machine Learning · Computer Science 2026-01-14 Chaoqun Fei , Huanjiang Liu , Tinglve Zhou , Yangyang Li , Tianyong Hao

Multi-task Vehicle Routing Problems (VRPs) aim to minimize routing costs while satisfying diverse constraints. Existing solvers typically adopt a unified reinforcement learning (RL) framework to learn generalizable patterns across tasks.…

Artificial Intelligence · Computer Science 2026-03-03 Shuangchun Gui , Suyu Liu , Xuehe Wang , Zhiguang Cao

Lane-changing (LC) is a challenging scenario for connected and automated vehicles (CAVs) because of the complex dynamics and high uncertainty of the traffic environment. This challenge can be handled by deep reinforcement learning (DRL)…

Robotics · Computer Science 2024-07-04 Xue Yao , Shengren Hou , Serge P. Hoogendoorn , Simeon C. Calvert

As modern AI workloads increasingly rely on heterogeneous accelerators, ensuring high-bandwidth and layout-flexible data movements between accelerator memories has become a pressing challenge. Direct Memory Access (DMA) engines promise high…

Hardware Architecture · Computer Science 2025-08-13 Fanchen Kong , Yunhao Deng , Xiaoling Yi , Ryan Antonio , Marian Verhelst

As both ML training and inference are increasingly distributed, parallelization techniques that shard (divide) ML model across GPUs of a distributed system, are often deployed. With such techniques, there is a high prevalence of…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-12 Shagnik Pal , Shaizeen Aga , Suchita Pati , Mahzabeen Islam , Lizy K. John

Graph Convolutional Networks (GCNs) are pivotal in extracting latent information from graph data across various domains, yet their acceleration on mainstream GPUs is challenged by workload imbalance and memory access irregularity. To…

Hardware Architecture · Computer Science 2023-08-24 Xi Xie , Hongwu Peng , Amit Hasan , Shaoyi Huang , Jiahui Zhao , Haowen Fang , Wei Zhang , Tong Geng , Omer Khan , Caiwen Ding

To fulfill the low latency requirements of today's applications, deployment of RDMA in datacenters has become prevalent over the recent years. However, the in-order delivery requirement of RDMAs prevents them from leveraging powerful…

Networking and Internet Architecture · Computer Science 2024-12-12 Sana Mahmood , Jinqi Lu , Soudeh Ghorbani

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications,…

Networking and Internet Architecture · Computer Science 2025-04-01 Weijian Zheng , Jack Kordas , Tyler J. Skluzacek , Raj Kettimuthu , Ian Foster

Facing the congestion challenges of mixed road networks comprising expressways and arterial road networks, traditional control solutions fall short. To effectively alleviate traffic congestion in mixed road networks, it is crucial to clear…

Systems and Control · Electrical Eng. & Systems 2024-05-13 Yunran Di , Haotian Shi , Weihua Zhang , Heng Ding , Xiaoyan Zheng , Bin Ran

The emergence of intelligent applications and recent advances in the fields of computing and networks are driving the development of computing and networks convergence (CNC) system. However, existing researches failed to achieve…

Networking and Internet Architecture · Computer Science 2024-02-06 Yujiao Hu , Qingmin Jia , Meng Shen , Renchao Xie , Tao Huang , F. Richard Yu

Heterogeneous computing platforms consisting of general purpose processors (GPPs) and graphics processing units (GPUs) have become commonplace in personal mobile devices and embedded systems. For years, programming of these platforms was…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-11-11 Jani Boutellier , Ilkka Hautala

Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to efficiently overlap computation and communication in machine…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-13 Suchita Pati , Shaizeen Aga , Mahzabeen Islam , Ryan Quach , Saleel Kudchadker , Mohamed Assem Ibrahim

Deep neural network (DNN) hardware (HW) accelerators have achieved great success in improving DNNs' performance and efficiency. One key reason is dataflow in executing a DNN layer, including on-chip data partitioning, computation…

Machine Learning · Computer Science 2024-10-10 Peng Xu , Wenqi Shao , Mingyu Ding , Ping Luo

On-device machine learning (ODML) enables powerful edge applications, but power consumption remains a key challenge for resource-constrained devices. To address this, developers often face a trade-off between model accuracy and power…

Machine Learning · Computer Science 2024-05-06 Haiguang Li , Usama Pervaiz , Joseph Antognini , Michał Matuszak , Lawrence Au , Gilles Roux , Trausti Thormundsson

Efficient management of traffic flow in urban environments presents a significant challenge, exacerbated by dynamic changes and the sheer volume of data generated by modern transportation networks. Traditional centralized traffic management…

Machine Learning · Computer Science 2025-01-29 Bob Johnson , Michael Geller

Generating molecular graphs is crucial in drug design and discovery but remains challenging due to the complex interdependencies between nodes and edges. While diffusion models have demonstrated their potentiality in molecular graph design,…

Machine Learning · Computer Science 2024-11-11 Xiaoyang Hou , Tian Zhu , Milong Ren , Dongbo Bu , Xin Gao , Chunming Zhang , Shiwei Sun

LLM training at the scale of tens of thousands of GPUs now spans multiple datacenters (DC), making cross-DC collectives over long-haul links unavoidable. A critical and overlooked bottleneck arises when these collectives collide with…

Networking and Internet Architecture · Computer Science 2026-05-14 Mariano Scazzariello , Noga H. Rotman , Dima Gavrilenko , Sajy Khashab , Alexander Shpiner , Matty Kadosh , Marco Chiesa , Dejan Kostic , Mark Silberstein

The exponential growth of data traffic and the increasing complexity of networked applications demand effective solutions capable of passively inspecting and analysing the network traffic for monitoring and security purposes. Implementing…

Networking and Internet Architecture · Computer Science 2024-07-24 Luca Deri , Alfredo Cardigliano , Francesco Fusco

The evolution from wired system to the wireless environment opens a set of challenge for the improvement of the wireless system performances because of many of their weakness compared to wired networks. To achieve this goal, cross layer…

Networking and Internet Architecture · Computer Science 2014-10-02 Mahmadou Issoufou Tiado , Riadh Dhaou , André-Luc Beylot
‹ Prev 1 4 5 6 7 8 10 Next ›