在三维环境中通过视觉目标预测将指令映射为动作
计算与语言
2019-03-19 v2
摘要
我们提出将指令执行分解为目标预测与动作生成。我们设计了一个模型,使用 LINGUNET(一种语言条件图像生成网络)将原始视觉观测映射到目标,然后生成完成这些目标所需的动作。我们的模型仅从演示中训练,无需外部资源。为评估我们的方法,我们引入了两个指令跟随基准:LANI,一项导航任务;以及 CHAI,其中智能体执行家庭指令。我们的评估展示了模型分解的优势,并阐明了新基准所带来的挑战。
引用
@article{arxiv.1809.00786,
title = {Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction},
author = {Dipendra Misra and Andrew Bennett and Valts Blukis and Eyvind Niklasson and Max Shatkhin and Yoav Artzi},
journal= {arXiv preprint arXiv:1809.00786},
year = {2019}
}
备注
Accepted at EMNLP 2018