English

End-to-End Throughput Benchmarking of Portable Deterministic CNN-Based Signal Processing Pipelines

Performance 2026-02-09 v1

Abstract

This paper presents a benchmarking methodology for evaluating end-to-end performance of deterministic signal-processing pipelines expressed using CNN-compatible primitives. The benchmark targets phased-array workloads such as ultrasound imaging and evaluates complete RF-to-image pipelines under realistic execution conditions. Performance is reported using sustained input throughput (MB/s), effective frame rate (FPS), and, where available, incremental energy per run and peak memory usage. Using this methodology, we benchmark a single deterministic, training-free CNN-based signal-processing pipeline executed unmodified across heterogeneous accelerator platforms, including an NVIDIA RTX 5090 GPU and a Google TPU v5e-1. The results demonstrate how different operator formulations (dynamic indexing, fully CNN-expressed, and sparse-matrix-based) impact performance and portability across architectures. This work is motivated by the need for portable, certifiable signal-processing implementations that avoid hardware-specific refactoring while retaining high performance on modern AI accelerators.

Keywords

Cite

@article{arxiv.2602.06216,
  title  = {End-to-End Throughput Benchmarking of Portable Deterministic CNN-Based Signal Processing Pipelines},
  author = {Christiaan Boerkamp and Akhil John Thomas},
  journal= {arXiv preprint arXiv:2602.06216},
  year   = {2026}
}

Comments

6 pages, 3 tables, Authors contributed equally

R2 v1 2026-07-01T10:23:26.513Z