English

Imitation Learning of Robot Policies by Combining Language, Vision and Demonstration

Robotics 2019-11-27 v1 Computation and Language Computer Vision and Pattern Recognition Machine Learning

Abstract

In this work we propose a novel end-to-end imitation learning approach which combines natural language, vision, and motion information to produce an abstract representation of a task, which in turn is used to synthesize specific motion controllers at run-time. This multimodal approach enables generalization to a wide variety of environmental conditions and allows an end-user to direct a robot policy through verbal communication. We empirically validate our approach with an extensive set of simulations and show that it achieves a high task success rate over a variety of conditions while remaining amenable to probabilistic interpretability.

Keywords

Cite

@article{arxiv.1911.11744,
  title  = {Imitation Learning of Robot Policies by Combining Language, Vision and Demonstration},
  author = {Simon Stepputtis and Joseph Campbell and Mariano Phielipp and Chitta Baral and Heni Ben Amor},
  journal= {arXiv preprint arXiv:1911.11744},
  year   = {2019}
}

Comments

Accepted to the NeurIPS 2019 Workshop on Robot Learning: Control and Interaction in the Real World, Vancouver, Canada