English

Deep Contextual Bandits for Orchestrating Multi-User MISO Systems with Multiple RISs

Information Theory 2022-02-17 v1 Signal Processing math.IT

Abstract

The emergent technology of Reconfigurable Intelligent Surfaces (RISs) has the potential to transform wireless environments into controllable systems, through programmable propagation of information-bearing signals. Techniques stemming from the field of Deep Reinforcement Learning (DRL) have recently gained popularity in maximizing the sum-rate performance in multi-user communication systems empowered by RISs. Such approaches are commonly based on Markov Decision Processes (MDPs). In this paper, we instead investigate the sum-rate design problem under the scope of the Multi-Armed Bandits (MAB) setting, which is a relaxation of the MDP framework. Nevertheless, in many cases, the MAB formulation is more appropriate to the channel and system models under the assumptions typically made in the RIS literature. To this end, we propose a simpler DRL approach for orchestrating multiple metasurfaces in RIS-empowered multi-user Multiple-Input Single-Output (MISO) systems, which we numerically show to perform equally well with a state-of-the-art MDP-based approach, while being less demanding computationally.

Keywords

Cite

@article{arxiv.2202.08194,
  title  = {Deep Contextual Bandits for Orchestrating Multi-User MISO Systems with Multiple RISs},
  author = {Kyriakos Stylianopoulos and George Alexandropoulos and Chongwen Huang and Chau Yuen and Mehdi Bennis and and Mérouane Debbah},
  journal= {arXiv preprint arXiv:2202.08194},
  year   = {2022}
}

Comments

6 pages, 4 figures, to be presented in IEEE ICC 2022

R2 v1 2026-06-24T09:41:19.049Z