English

Reinforcement and Imitation Learning via Interactive No-Regret Learning

Machine Learning 2014-06-24 v1 Machine Learning

Abstract

Recent work has demonstrated that problems-- particularly imitation learning and structured prediction-- where a learner's predictions influence the input-distribution it is tested on can be naturally addressed by an interactive approach and analyzed using no-regret online learning. These approaches to imitation learning, however, neither require nor benefit from information about the cost of actions. We extend existing results in two directions: first, we develop an interactive imitation learning approach that leverages cost information; second, we extend the technique to address reinforcement learning. The results provide theoretical support to the commonly observed successes of online approximate policy iteration. Our approach suggests a broad new family of algorithms and provides a unifying view of existing techniques for imitation and reinforcement learning.

Keywords

Cite

@article{arxiv.1406.5979,
  title  = {Reinforcement and Imitation Learning via Interactive No-Regret Learning},
  author = {Stephane Ross and J. Andrew Bagnell},
  journal= {arXiv preprint arXiv:1406.5979},
  year   = {2014}
}

Comments

14 pages. Under review for NIPS 2014 conference

R2 v1 2026-06-22T04:45:00.690Z