The ability to learn and execute optimal control policies safely is critical to realization of complex autonomy, especially where task restarts are not available and/or the systems are safety-critical. Safety requirements are often expressed in terms of state and/or control constraints. Methods such as barrier transformation and control barrier functions have been successfully used, in conjunction with model-based reinforcement learning, for safe learning in systems under state constraints, to learn the optimal control policy. However, existing barrier-based safe learning methods rely on full state feedback. In this paper, an output-feedback safe model-based reinforcement learning technique is developed that utilizes a novel dynamic state estimator to implement simultaneous learning and control for a class of safety-critical systems with partially observable state.
@article{arxiv.2110.00271,
title = {Safety aware model-based reinforcement learning for optimal control of a class of output-feedback nonlinear systems},
author = {S M Nahid Mahmud and Moad Abudia and Scott A Nivison and Zachary I. Bell and Rushikesh Kamalapurkar},
journal= {arXiv preprint arXiv:2110.00271},
year = {2021}
}
Comments
arXiv admin note: substantial text overlap with arXiv:2007.12666