English

Solving infinite-horizon POMDPs with memoryless stochastic policies in state-action space

Machine Learning 2022-05-30 v1 Systems and Control Systems and Control Optimization and Control

Abstract

Reward optimization in fully observable Markov decision processes is equivalent to a linear program over the polytope of state-action frequencies. Taking a similar perspective in the case of partially observable Markov decision processes with memoryless stochastic policies, the problem was recently formulated as the optimization of a linear objective subject to polynomial constraints. Based on this we present an approach for Reward Optimization in State-Action space (ROSA). We test this approach experimentally in maze navigation tasks. We find that ROSA is computationally efficient and can yield stability improvements over other existing methods.

Keywords

Cite

@article{arxiv.2205.14098,
  title  = {Solving infinite-horizon POMDPs with memoryless stochastic policies in state-action space},
  author = {Johannes Müller and Guido Montúfar},
  journal= {arXiv preprint arXiv:2205.14098},
  year   = {2022}
}

Comments

Accepted as an extended abstract at RLDM 2022, 5 pages, 2 figures

R2 v1 2026-06-24T11:31:11.547Z