VL-Fields:面向语言接地的神经隐式空间表示
计算机视觉与模式识别
2023-05-26 v2
摘要
我们提出视觉-语言场(VL-Fields),一种支持开放词汇语义查询的神经隐式空间表示。我们的模型通过从语言驱动的分割模型中蒸馏信息,将场景几何与视觉-语言训练的潜在特征进行编码与融合。VL-Fields 在训练时无需场景物体类别的先验知识,这使其成为机器人领域一种有前景的表示。我们的模型在语义分割任务上以近 10% 的优势优于类似的 CLIP-Fields 模型。
引用
@article{arxiv.2305.12427,
title = {VL-Fields: Towards Language-Grounded Neural Implicit Spatial Representations},
author = {Nikolaos Tsagkas and Oisin Mac Aodha and Chris Xiaoxuan Lu},
journal= {arXiv preprint arXiv:2305.12427},
year = {2023}
}
备注
Project page: https://tsagkas.github.io/vl-fields/