English
Related papers

Related papers: The DMA Streaming Framework: Kernel-Level Buffer O…

200 papers

The dispersed node locations and complex topologies of edge networks, combined with intricate dynamic microservice dependencies, render traditional centralized microservice architectures (MSAs) unsuitable. In this paper, we propose a…

Networking and Internet Architecture · Computer Science 2025-01-03 Yuang Chen , Chengdi Lu , Yongsheng Huang , Chang Wu , Fengqian Guo , Hancheng Lu , Chang Wen Chen

Next-generation wireless networks will rely on mmWave/sub-THz spectrum and extremely large antenna arrays (ELAAs). This will push their operation into the near field where far-field beam management degrades and beam training becomes more…

Signal Processing · Electrical Eng. & Systems 2025-11-18 Zijun Wang , Anjali Omer , Jacob Chakareski , Nicholas Mastronarde , Rui Zhang

Faced with the massive connection, sporadic transmission, and small-sized data packets in future cellular communication, a grant-free non-orthogonal random access (NORA) system is considered in this paper, which could reduce the access…

Information Theory · Computer Science 2019-10-10 Zhaoji Zhang , Ying Li , Chongwen Huang , Qinghua Guo , Chau Yuen , Yong Liang Guan

High-speed, low-latency obstacle avoidance that is insensitive to sensor noise is essential for enabling multiple decentralized robots to function reliably in cluttered and dynamic environments. While other distributed multi-agent collision…

Artificial Intelligence · Computer Science 2017-07-07 Pinxin Long , Wenxi Liu , Jia Pan

Massive numbers of nodes will be connected in future wireless networks. This brings great difficulty to collect a large amount of data. Instead of collecting the data individually, computation over multi-access channel (CoMAC) provides an…

Information Theory · Computer Science 2018-12-14 Fangzhou Wu , Li Chen , Nan Zhao , Yunfei Chen , F. Richard Yu , Guo Wei

High performance large scale graph analytics are essential to timely analyze relationships in big data sets. Conventional processor architectures suffer from inefficient resource usage and bad scaling on those workloads. To enable efficient…

Denoising Diffusion Probabilistic Models (DDPMs) have established a new state-of-the-art in generative image synthesis, yet their deployment is hindered by significant computational overhead during inference, often requiring up to 1,000…

Machine Learning · Computer Science 2025-11-25 Srishti Gupta , Yashasvee Taiwade

As the artificial intelligence community advances into the era of large models with billions of parameters, distributed training and inference have become essential. While various parallelism strategies-data, model, sequence, and…

Machine Learning · Computer Science 2025-03-13 Ruifeng She , Bowen Pang , Kai Li , Zehua Liu , Tao Zhong

Novel non-volatile memory (NVM) technologies offer high-speed and high-density data storage. In addition, they overcome the von Neumann bottleneck by enabling computing-in-memory (CIM). Various computer architectures have been proposed to…

Cryptography and Security · Computer Science 2023-04-13 Lennart M. Reimann , Felix Staudigl , Rainer Leupers

Network-wide traffic analytics are often needed for various network monitoring tasks. These measurements are often performed by collecting samples at network switches, which are then sent to the controller for aggregation. However,…

Networking and Internet Architecture · Computer Science 2020-05-27 Ran Ben Basat , Xiaoqi Chen , Gil Einziger , Shir Landau Feibish , Danny Raz , Minlan Yu

Many reasons make NFV an attractive paradigm for IT security: lowers costs, agile operations and better isolation as well as fast security updates, improved incident responses and better level of automation. At the same time, the network…

Networking and Internet Architecture · Computer Science 2016-08-05 Pier Luigi Ventre , Alberto Caponi , Davide Palmisano , Stefano Salsano , Giuseppe Siracusano , Marco Bonola , Giuseppe Bianchi

In this paper we propose Virtuoso, a purely software-based multi-path RDMA solution for data center networks (DCNs) to effectively utilize the rich multi-path topology for load balancing and reliability. As a "middleware" library operating…

Networking and Internet Architecture · Computer Science 2020-09-02 Feng Tian , Wendi Feng , Yang Zhang , Zhi-Li Zhang

Spiking neural networks excel at event-driven sensing. Yet, maintaining task-relevant context over long timescales both algorithmically and in hardware, while respecting both tight energy and memory budgets, remains a core challenge in the…

Neural and Evolutionary Computing · Computer Science 2026-05-05 Pengfei Sun , Zhe Su , Jascha Achterberg , Giacomo Indiveri , Dan F. M. Goodman , Danyal Akarca

Limited memory bandwidth is a critical bottleneck in modern systems. 3D-stacked DRAM enables higher bandwidth by leveraging wider Through-Silicon-Via (TSV) channels, but today's systems cannot fully exploit them due to the limited internal…

Hardware Architecture · Computer Science 2015-06-11 Donghyuk Lee , Gennady Pekhimenko , Samira Khan , Saugata Ghose , Onur Mutlu

Cutting-edge embedded system applications, such as self-driving cars and unmanned drone software, are reliant on integrated CPU/GPU platforms for their DNNs-driven workload, such as perception and other highly parallel components. In this…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-03-20 Soroush Bateni , Zhendong Wang , Yuankun Zhu , Yang Hu , Cong Liu

High-speed RDMA networks are getting rapidly adopted in the industry for their low latency and reduced CPU overheads. To verify that RDMA can be used in production, system administrators need to understand the set of application workloads…

Networking and Internet Architecture · Computer Science 2023-04-25 Xinhao Kong , Yibo Zhu , Huaping Zhou , Zhuo Jiang , Jianxi Ye , Chuanxiong Guo , Danyang Zhuo

Processing in-memory (PIM) is promising to accelerate neural networks (NNs) because it minimizes data movement and provides large computational parallelism. Similar to machine learning accelerators, application mapping, which determines the…

Hardware Architecture · Computer Science 2024-07-02 Xuan Wang , Minxuan Zhou , Tajana Rosing

Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to efficiently overlap computation and communication in machine…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-13 Suchita Pati , Shaizeen Aga , Mahzabeen Islam , Ryan Quach , Saleel Kudchadker , Mohamed Assem Ibrahim

This paper is focused on code-domain non-orthogonal multiple access (CD-NOMA), which is an emerging paradigm to support massive connectivity for future machine-type wireless networks. We take a comparative approach to study two types of…

Information Theory · Computer Science 2020-09-10 Zilong Liu , Lie-Liang Yang

In this paper, we present a framework for moving compute and data between processing elements in a distributed heterogeneous system. The implementation of the framework is based on the LLVM compiler toolchain combined with the UCX…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-12-13 Wenbin Lu , Luis E. Peña , Pavel Shamis , Valentin Churavy , Barbara Chapman , Steve Poole