English
Related papers

Related papers: A Lightweight High-Throughput Collective-Capable N…

200 papers

GPUs are the heart of the latest generations of supercomputers. We efficiently accelerate a compressible multiphase flow solver via OpenACC on NVIDIA and AMD Instinct GPUs. Optimization is accomplished by specifying the directive clauses…

This paper proposes and optimizes a cooperative non-orthogonal multiple-access (Co-NOMA) scheme in the context of multicell visible light communications (VLC) networks, as a means to mitigate inter-cell interference in Co-NOMA-enabled…

Information Theory · Computer Science 2020-05-20 Mohanad Obeed , Hayssam Dahrouj , Anas M. Salhab , Anas Chaaban , Salam A. Zummo , Mohammed-Slim Alouini

Non-orthogonal multiple access (NOMA) is a key technology to enable massive machine type communications (mMTC) in 5G networks and beyond. In this paper, NOMA is applied to improve the random access efficiency in high-density…

Networking and Internet Architecture · Computer Science 2024-10-28 Sami Khairy , Prasanna Balaprakash , Lin X. Cai , H. Vincent Poor

The acceleration of pruned Deep Neural Networks (DNNs) on edge devices such as Microcontrollers (MCUs) is a challenging task, given the tight area- and power-constraints of these devices. In this work, we propose a three-fold contribution…

Machine Learning · Computer Science 2025-03-20 Francesco Daghero , Daniele Jahier Pagliari , Francesco Conti , Luca Benini , Massimo Poncino , Alessio Burrello

Channel state information (CSI) at transmitter is crucial for massive MIMO downlink systems to achieve high spectrum and energy efficiency. Existing works have provided deep learning architectures for CSI feedback and recovery at the…

Signal Processing · Electrical Eng. & Systems 2022-04-21 Yu-Chien Lin , Ta-Sung Lee , Zhi Ding

The rapid adoption of large language models (LLMs) is pushing AI accelerators toward increasingly powerful and specialized designs. Instead of further complicating software development with deeply hierarchical scratchpad memories (SPMs) and…

Hardware Architecture · Computer Science 2025-12-09 Zhongchun Zhou , Chengtao Lai , Yuhang Gu , Wei Zhang

Recently, high-speed and short-distance networks are widely deployed and their necessity is rapidly increasing everyday. This type of networks is used in several network applications; such as Local Area Networks (LAN) and Data Center…

Networking and Internet Architecture · Computer Science 2016-01-25 Mohamed A. Alrshah , Mohamed Othman , Borhanuddin Ali , Zurina Mohd Hanapi

Modern Deep Neural Network (DNN) accelerators are equipped with increasingly larger on-chip buffers to provide more opportunities to alleviate the increasingly severe DRAM bandwidth pressure. However, most existing research on buffer…

Hardware Architecture · Computer Science 2025-01-23 Jingwei Cai , Xuan Wang , Mingyu Gao , Sen Peng , Zijian Zhu , Yuchen Wei , Zuotong Wu , Kaisheng Ma

On-device large language models commonly employ task-specific adapters (e.g., LoRAs) to deliver strong performance on downstream tasks. While storing all available adapters is impractical due to memory constraints, mobile devices typically…

Machine Learning · Computer Science 2026-01-27 Ondrej Bohdal , Taha Ceritli , Mete Ozay , Jijoong Moon , Kyeng-Hun Lee , Hyeonmok Ko , Umberto Michieli

Emerging ReRAM-based accelerators process neural networks via analog Computing-in-Memory (CiM) for ultra-high energy efficiency. However, significant overhead in peripheral circuits and complex nonlinear activation modes constrain system…

Hardware Architecture · Computer Science 2024-12-31 Peng Dang , Huawei Li , Wei Wang

Next-generation wireless technologies (for immersive-massive communication, joint communication and sensing) demand highly parallel architectures for massive data processing. A common architectural template scales up by grouping tens to…

Hardware Architecture · Computer Science 2025-07-08 Samuel Riedel , Yichao Zhang , Marco Bertuletti , Luca Benini

Mobile edge computing (MEC) networks bring computing and storage capabilities closer to edge devices, which reduces latency and improves network performance. However, to further reduce transmission and computation costs while satisfying…

Information Theory · Computer Science 2023-09-28 Xiangyu Gao , Yaping Sun , Hao Chen , Xiaodong Xu , Shuguang Cui

Multiple-input multiple-output non-orthogonal multiple access (MIMO-NOMA) cellular network is promising for supporting massive connectivity. This paper exploits low-latency machine learning in the MIMO-NOMA uplink transmission environment,…

Information Theory · Computer Science 2021-06-29 Mian Guo , Chun Shan , Mithun Mukherjee , Jaime Lloret , Quansheng Guan

Efficient radio spectrum utilization and low energy consumption in mobile devices are essential in developing next generation wireless networks. This paper presents a new medium access control (MAC) mechanism to enhance spectrum efficiency…

Networking and Internet Architecture · Computer Science 2016-11-15 Kamal Rahimi Malekshan , Weihua Zhuang , Yves Lostanlen

Collective communication is becoming increasingly important in data center and supercomputer workloads with an increase in distributed AI related jobs. However, existing libraries that provide collective support such as NCCL, RCCL, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-17 Siddharth Singh , Keshav Pradeep , Mahua Singh , Cunyang Wei , Abhinav Bhatele

Future wireless networks are expected to connect large-scale low-powered communication devices using the available spectrum resources. Backscatter communications (BC) is an emerging technology towards battery-free transmission in future…

Signal Processing · Electrical Eng. & Systems 2022-08-03 Wali Ullah Khan , Eva Lagunas , Asad Mahmood , Zain Ali , Symeon Chatzinotas , Björn Ottersten , Octavia A. Dobre

To support Machine Type Communications (MTC) in next generation mobile networks, NarrowBand-IoT (NB-IoT) has been released by the Third Generation Partnership Project (3GPP) as a promising solution to provide extended coverage and low…

Networking and Internet Architecture · Computer Science 2018-12-24 Ali Shahini , Nirwan Ansari

The ever-increasing computation complexity of fast-growing Deep Neural Networks (DNNs) has requested new computing paradigms to overcome the memory wall in conventional Von Neumann computing architectures. The emerging Computing-In-Memory…

Hardware Architecture · Computer Science 2021-07-21 Kaining Zhou , Yangshuo He , Rui Xiao , Kejie Huang

DRAM-based in-situ accelerators have shown their potential in addressing the memory wall challenge of the traditional von Neumann architecture. Such accelerators exploit charge sharing or logic circuits for simple logic operations at the…

Hardware Architecture · Computer Science 2024-02-08 Minki Jeong , Wanyeong Jung

Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resource underutilization and off-chip communication overheads.…

Hardware Architecture · Computer Science 2026-02-17 Zongle Huang , Hongyang Jia , Kaiwei Zou , Yongpan Liu