English

Few-shot Object Grounding and Mapping for Natural Language Robot Instruction Following

Robotics 2020-11-17 v1 Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Machine Learning

Abstract

We study the problem of learning a robot policy to follow natural language instructions that can be easily extended to reason about new objects. We introduce a few-shot language-conditioned object grounding method trained from augmented reality data that uses exemplars to identify objects and align them to their mentions in instructions. We present a learned map representation that encodes object locations and their instructed use, and construct it from our few-shot grounding output. We integrate this mapping approach into an instruction-following policy, thereby allowing it to reason about previously unseen objects at test-time by simply adding exemplars. We evaluate on the task of learning to map raw observations and instructions to continuous control of a physical quadcopter. Our approach significantly outperforms the prior state of the art in the presence of new objects, even when the prior approach observes all objects during training.

Keywords

Cite

@article{arxiv.2011.07384,
  title  = {Few-shot Object Grounding and Mapping for Natural Language Robot Instruction Following},
  author = {Valts Blukis and Ross A. Knepper and Yoav Artzi},
  journal= {arXiv preprint arXiv:2011.07384},
  year   = {2020}
}

Comments

4th Conference on Robot Learning (CoRL 2020), Cambridge MA, USA

R2 v1 2026-06-23T20:13:26.728Z