English

Real-Time and Scalable Zak-OTFS Receiver Processing on GPUs

Signal Processing 2026-04-03 v1 Networking and Internet Architecture

Abstract

Orthogonal time frequency space (OTFS) modulation offers superior robustness to high-mobility channels compared to conventional orthogonal frequency-division multiplexing (OFDM) waveforms. However, its explicit delay-Doppler (DD) domain representation incurs substantial signal processing complexity, especially with increased DD domain grid sizes. To address this challenge, we present a scalable, real-time Zak-OTFS receiver architecture on GPUs through hardware--algorithm co-design that exploits DD-domain channel sparsity. Our design leverages compact matrix operations for key processing stages, a branchless iterative equalizer, and a structured sparse channel matrix of the DD domain channel matrix to significantly reduce computational and memory overhead. These optimizations enable low-latency processing that consistently meets the 99.9-th percentile real-time processing deadline. The proposed system achieves up to 906.52 Mbps throughput with a DD grid size of (16384,32) using 16QAM modulation over 245.76 MHz bandwidth. Extensive evaluations under a Vehicular-A channel model demonstrate strong scalability and robust performance across CPU (Intel Xeon) and multiple GPU platforms (NVIDIA Jetson Orin, RTX 6000 Ada, A100, and H200), highlighting the effectiveness of compute-aware Zak-OTFS receiver design for next-generation (NextG) high-mobility communication systems.

Keywords

Cite

@article{arxiv.2604.02266,
  title  = {Real-Time and Scalable Zak-OTFS Receiver Processing on GPUs},
  author = {Junyao Zheng and Chung-Hsuan Tung and Yuncheng Yao and Nishant Mehrotra and Sandesh Mattu and Zhenzhou Qi and Danyang Zhuo and Robert Calderbank and Tingjun Chen},
  journal= {arXiv preprint arXiv:2604.02266},
  year   = {2026}
}

Comments

This work has been submitted to the IEEE for possible publication

R2 v1 2026-07-01T11:51:31.240Z