基于随机策略梯度的语义通信无模型强化学习
信号处理
2026-02-18 v2 信息论
机器学习
math.IT
机器学习
摘要
继机器学习工具在无线通信中近期取得成功之后,Weaver 于 1949 年提出的语义通信思想已受到关注。它打破了 Shannon 的经典设计范式,旨在传输消息的语义(即含义)而非其精确版本,从而节省信息率。在本工作中,我们应用随机策略梯度(Stochastic Policy Gradient, SPG)通过强化学习设计语义通信系统,将发送端与接收端分离,且不需要已知或可微的信道模型——这是朝实际部署迈进的关键一步。此外,我们从接收变量与目标变量间互信息最大化的角度,推导了 SPG 在经典通信与语义通信中的使用。数值结果表明,尽管收敛速度有所下降,我们的方法与基于重参数化技巧的模型感知方法取得了相当的性能。
引用
@article{arxiv.2305.03571,
title = {Model-free Reinforcement Learning of Semantic Communication by Stochastic Policy Gradient},
author = {Edgar Beck and Carsten Bockelmann and Armin Dekorsy},
journal= {arXiv preprint arXiv:2305.03571},
year = {2026}
}
备注
Accepted for publication in IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN 2024), Source Code: https://github.com/ant-uni-bremen/SINFONY