English

Symbol Distributions in Semantic Communications: A Source-Channel Equilibrium Perspective

Information Theory 2025-12-17 v1 Signal Processing math.IT

Abstract

Semantic communication systems often use an end-to-end neural network to map input data into continuous symbols. These symbols, which are essentially neural network features, usually have fixed dimensions and heavy-tailed distributions. However, due to the end-to-end training nature of the neural network encoder, the underlying reason for the symbol distribution remains underexplored. We propose a new explanation for the semantic symbol distribution: an inherent trade-off between source coding and communications. Specifically, the encoder balances two objectives: allocating power for minimum \emph{effective codelength} (for source coding) and maximizing mutual information (for communications). We formalize this trade-off via an information-theoretic optimization framework, which yields a Student's tt-distribution as the resulting symbol distribution. Through extensive studies on image-based semantic systems, we find that our formulation models the learned symbols and predicts how the symbol distribution's shape parameter changes with respect to (i) the use of variable-length coding and (ii) the dataset's entropy variability. Furthermore, we demonstrate how introducing a regularizer that enforces a target symbol distribution, which guides the encoder towards a target prior (e.g., Gaussian), improves training convergence and supports our hypothesis.

Keywords

Cite

@article{arxiv.2512.14022,
  title  = {Symbol Distributions in Semantic Communications: A Source-Channel Equilibrium Perspective},
  author = {Hanju Yoo and Dongha Choi and Songkuk Kim and Chan-Byoung Chae and Robert W. Heath},
  journal= {arXiv preprint arXiv:2512.14022},
  year   = {2025}
}
R2 v1 2026-07-01T08:26:32.196Z