English

On the connection between Bregman divergence and value in regularized Markov decision processes

Machine Learning 2022-11-08 v4 Artificial Intelligence Optimization and Control

Abstract

In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a regularized Markov decision process. This result has implications for multi-task reinforcement learning, offline reinforcement learning, and regret analysis under function approximation, among others.

Cite

@article{arxiv.2210.12160,
  title  = {On the connection between Bregman divergence and value in regularized Markov decision processes},
  author = {Brendan O'Donoghue},
  journal= {arXiv preprint arXiv:2210.12160},
  year   = {2022}
}
R2 v1 2026-06-28T04:12:32.084Z