从大语言模型中提取零样本常识用于机器人三维场景理解
机器人学
2022-06-22 v2 计算与语言
摘要
语义三维场景理解是机器人领域中至关重要的问题。尽管同步定位与建图算法已取得显著进展,机器人仍远未具备普通人类关于家用物体及其位置的常识知识。我们引入一种新方法,利用大语言模型(LLM)中嵌入的常识,根据所含物体对房间进行标注。该算法具有额外优势:(i)无需任务特定预训练(完全在零样本机制下运行);(ii)可泛化至任意房间与物体标签,包括此前未见的标签——二者均为机器人场景理解算法中极受青睐的特性。所提算法作用于现代空间感知系统生成的 3D 场景图,我们希望其能为机器人更可泛化、可扩展的高层三维场景理解铺平道路。
引用
@article{arxiv.2206.04585,
title = {Extracting Zero-shot Common Sense from Large Language Models for Robot 3D Scene Understanding},
author = {William Chen and Siyi Hu and Rajat Talak and Luca Carlone},
journal= {arXiv preprint arXiv:2206.04585},
year = {2022}
}
备注
4 pages (excluding references and appendix), 2 figures, 2 tables. Submitted to Robotics: Science and Systems 2022 2nd Workshop on Scaling Robot Learning. Corrected typos and notation