Ges3ViG: 将指向手势融入基于语言的 3D 视觉定位
计算机视觉与模式识别
2025-04-15 v1 人工智能
多媒体
摘要
3D 躲体参考理解 (3D-ERU) 结合语言描述与伴随指向手势,以识别 3D 场景中最相关的目标对象。尽管已有工作探索了纯语言-based 的 3D 定位,但针对融合人类指向手势的 3D-ERU 研究仍有限。为填补这一空白,我们引入了数据增强框架 Imputer,并利用其为 3D-ERU 构建新基准数据集 ImputeRefer。我们还提出了 Ges3ViG,这一一种 novel 3D-ERU 模型,实验表明其在各种 3D-ERU 模型中准确率提升约 30%,在纯语言-based 的 3D 定位模型中提升约 9%。我们的代码与数据集已公开于 https://github.com/AtharvMane/Ges3ViG。
引用
@article{arxiv.2504.09623,
title = {Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding},
author = {Atharv Mahesh Mane and Dulanga Weerakoon and Vigneshwaran Subbaraju and Sougata Sen and Sanjay E. Sarma and Archan Misra},
journal= {arXiv preprint arXiv:2504.09623},
year = {2025}
}
备注
Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025