English
Related papers

Related papers: Parallel Interleaver Design for a High Throughput …

200 papers

Fixed-complexity Sphere Decoder (FSD) is a recently proposed technique for Multiple-Input Multiple-Output (MIMO) detection. It has several outstanding features such as constant throughput and large potential parallelism, which makes it…

Information Theory · Computer Science 2010-06-22 Bin Wu , Guido Masera

Neural network-based decoding methods show promise in enhancing error correction performance but face challenges with punctured codes. In particular, existing methods struggle to adapt to variable code rates or meet protocol compatibility…

Machine Learning · Computer Science 2025-10-31 Yongli Yan , Linglong Dai

Diffusion large language models (dLLMs) have recently drawn considerable attention within the research community as a promising alternative to autoregressive generation, offering parallel token prediction and lower inference latency. Yet,…

Computation and Language · Computer Science 2025-10-01 Zigeng Chen , Gongfan Fang , Xinyin Ma , Ruonan Yu , Xinchao Wang

Extrapolating ultra-long contexts (text length >128K) remains a major challenge for large language models (LLMs), as most training-free extrapolation methods are not only severely limited by memory bottlenecks, but also suffer from the…

Computation and Language · Computer Science 2025-06-10 Jing Xiong , Jianghan Shen , Chuanyang Zheng , Zhongwei Wan , Chenyang Zhao , Chiwun Yang , Fanghua Ye , Hongxia Yang , Lingpeng Kong , Ngai Wong

A critical aspect of reliable communication involves the design of codes that allow transmissions to be robustly and computationally efficiently decoded under noisy conditions. Advances in the design of reliable codes have been driven by…

Information Theory · Computer Science 2021-11-23 Karl Chahine , Yihan Jiang , Pooja Nuti , Hyeji Kim , Joonyoung Cho

Distributed inference of large language models (LLMs) using tensor parallelism can introduce communication overheads of $20$% even over GPUs connected via NVLink, a high-speed GPU interconnect. Several techniques have been proposed to…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-04 Raja Gond , Nipun Kwatra , Ramachandran Ramjee

In this paper, a distributed turbo-like coding scheme for wireless networks with relays is proposed. We consider a scenario where multiple sources communicate with a single destination with the help of a relay. The proposed scheme can be…

Information Theory · Computer Science 2009-10-13 Roua Youssef , Alexandre Graell i Amat

Polar codes are the first class of capacity-achieving forward error correction (FEC) codes. They have been selected as one of the coding schemes for the 5G communication systems due to their excellent error correction performance when…

Signal Processing · Electrical Eng. & Systems 2019-05-23 ChenYang Xia , YouZhe Fan , Chi-ying Tsui

The purpose of this study is to construct a near capacity Irregular Turbo Code and to evaluate its performance over Gaussian channel. The methodology used to evaluate and measure the performance of the new design is by simulating the system…

Information Theory · Computer Science 2016-04-06 Abiodun Sholiyi , Jafar A. Alzubi , Omar A. Alzubi , Omar Almomani , Tim O'Farrell

Achieving high image quality is an important aspect in an increasing number of wireless multimedia applications. These applications require resource efficient error correction hardware to detect and correct errors introduced by the…

Information Theory · Computer Science 2013-08-02 Vikram Arkalgud Chandrasetty , Syed Mahfuzul Aziz

Polar codes have received increasing attention in the past decade, and have been selected for the next generation of wireless communication standard. Most research on polar codes has focused on codes constructed from a $2\times2$…

Hardware Architecture · Computer Science 2018-02-05 Gabriele Coppolino , Carlo Condo , Guido Masera , Warren J. Gross

This paper summarizes the design of a programmable processor with transport triggered architecture (TTA) for decoding LDPC and turbo codes. The processor architecture is designed in such a manner that it can be programmed for LDPC or turbo…

Information Theory · Computer Science 2015-02-03 Shahriar Shahabuddin , Janne Janhunen , Muhammet Fatih Bayramoglu , Markku Juntti , Amanullah Ghazi , Olli Silven

Large deep learning models have demonstrated strong ability to solve many tasks across a wide range of applications. Those large models typically require training and inference to be distributed. Tensor parallelism is a common technique…

Large language model inference is both memory-intensive and time-consuming, often requiring distributed algorithms to efficiently scale. Various model parallelism strategies are used in multi-gpu training and inference to partition…

High-speed packet processing on multicore CPUs places extreme demands on memory allocators. In systems like DPDK, fixed-size memory pools back packet buffers (mbufs) to avoid costly dynamic allocation. However, even DPDK's optimized mempool…

Performance · Computer Science 2026-03-24 Junyi Yang

Ultra-reliable low latency communication (URLLC) is a key part of 5G wireless systems. Achieving low latency necessitates codes with short blocklengths for which polar codes with successive cancellation list (SCL) decoding typically…

Hardware Architecture · Computer Science 2025-12-22 Darja Nonaca , Jérémy Guichemerre , Reinhard Wiesmayr , Nihat Engin Tunali , Christoph Studer

In this paper, we propose a parallel block-based Viterbi decoder (PBVD) on the graphic processing unit (GPU) platform for the decoding of convolutional codes. The decoding procedure is simplified and parallelized, and the characteristic of…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-08-02 Hao Peng , Rongke Liu , Yi Hou , Ling Zhao

Offloading compute-intensive kernels to hardware accelerators relies on the large degree of parallelism offered by these platforms. However, the effective bandwidth of the memory interface often causes a bottleneck, hindering the…

Hardware Architecture · Computer Science 2022-02-25 Corentin Ferry , Tomofumi Yuki , Steven Derrien , Sanjay Rajopadhye

Recently, a parallel decoding framework of $G_N$-coset codes was proposed. High throughput is achieved by decoding the independent component polar codes in parallel. Various algorithms can be employed to decode these component codes,…

Information Theory · Computer Science 2020-04-22 Xianbin Wang , Jiajie Tong , Huazi Zhang , Shengchen Dai , Rong Li , Jun Wang

Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-28 Man Liu , Xingchen Liu , Xingjian Tian , Bing Lu , Shengkay Lyu , Shengquan Yin , Wenjing Huang , Zheng Wei , Hairui Zhao , Guangming Tan , Dingwen Tao