English

Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks

Machine Learning 2017-03-09 v3 Machine Learning

Abstract

We present an algorithm for model-based reinforcement learning that combines Bayesian neural networks (BNNs) with random roll-outs and stochastic optimization for policy learning. The BNNs are trained by minimizing α\alpha-divergences, allowing us to capture complicated statistical patterns in the transition dynamics, e.g. multi-modality and heteroskedasticity, which are usually missed by other common modeling approaches. We illustrate the performance of our method by solving a challenging benchmark where model-based approaches usually fail and by obtaining promising results in a real-world scenario for controlling a gas turbine.

Keywords

Cite

@article{arxiv.1605.07127,
  title  = {Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks},
  author = {Stefan Depeweg and José Miguel Hernández-Lobato and Finale Doshi-Velez and Steffen Udluft},
  journal= {arXiv preprint arXiv:1605.07127},
  year   = {2017}
}
R2 v1 2026-06-22T14:07:30.187Z