中文

基于开放词汇 3D 场景图的上下文感知实体定位

机器人学 2023-09-29 v1 计算机视觉与模式识别

摘要

我们提出开放词汇 3D 场景图(OVSG),一种用基于自由文本查询对各类实体(如物体实例、智能体和区域)进行定位的形式化框架。与传统的基于语义的物体定位方法不同,我们的系统支持上下文感知的实体定位,允许诸如“拿起厨房桌上的一只杯子”或“导航到有人坐着的沙发”之类的查询。与现有关于 3D 场景图的研究相比,OVSG 支持自由文本输入与开放词汇查询。通过使用 ScanNet 数据集和自采集数据集的一系列对比实验,我们证明所提方法显著超越先前基于语义的定位技术的性能。此外,我们展示了 OVSG 在真实世界机器人导航与操作实验中的实际应用。

关键词

引用

@article{arxiv.2309.15940,
  title  = {Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs},
  author = {Haonan Chang and Kowndinya Boyalakuntla and Shiyang Lu and Siwei Cai and Eric Jing and Shreesh Keskar and Shijie Geng and Adeeb Abbas and Lifeng Zhou and Kostas Bekris and Abdeslam Boularias},
  journal= {arXiv preprint arXiv:2309.15940},
  year   = {2023}
}

备注

The code and dataset used for evaluation can be found at https://github.com/changhaonan/OVSG}{https://github.com/changhaonan/OVSG. This paper has been accepted by CoRL2023