English
Related papers

Related papers: LidaRefer: Context-aware Outdoor 3D Visual Groundi…

200 papers

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Zhenxiang Lin , Xidong Peng , Peishan Cong , Ge Zheng , Yujin Sun , Yuenan Hou , Xinge Zhu , Sibei Yang , Yuexin Ma

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Xinyi Wang , Na Zhao , Zhiyuan Han , Dan Guo , Xun Yang

3D visual grounding aims to localize the object in 3D point cloud scenes that semantically corresponds to given natural language sentences. It is very critical for roadside infrastructure system to interpret natural languages and localize…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Panquan Yang , Junfei Huang , Zongzhangbao Yin , Yingsong Hu , Anni Xu , Xinyi Luo , Xueqi Sun , Hai Wu , Sheng Ao , Zhaoxing Zhu , Chenglu Wen , Cheng Wang

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar objects through the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Feng Xiao , Hongbin Xu , Guocan Zhao , Wenxiong Kang

3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. Previous works usually require significant data relating to point color and their descriptions to exploit…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Tung-Yu Wu , Sheng-Yu Huang , Yu-Chiang Frank Wang

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Zoey Guo , Yiwen Tang , Ray Zhang , Dong Wang , Zhigang Wang , Bin Zhao , Xuelong Li

This paper addresses the problem of 3D referring expression comprehension (REC) in autonomous driving scenario, which aims to ground a natural language to the targeted region in LiDAR point clouds. Previous approaches for REC usually focus…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Wenhao Cheng , Junbo Yin , Wei Li , Ruigang Yang , Jianbing Shen

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

Robotics · Computer Science 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

We present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent tendencies of overfitting a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Motoki Kawanabe

3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications, which involve searching for objects or patterns using…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Austin T. Wang , ZeMing Gong , Angel X. Chang

Recent progress in 3D scene understanding has explored visual grounding (3DVG) to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Zhihao Yuan , Xu Yan , Zhuo Li , Xuhao Li , Yao Guo , Shuguang Cui , Zhen Li

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding…

Computer Vision and Pattern Recognition · Computer Science 2021-07-30 Zhihao Yuan , Xu Yan , Yinghong Liao , Ruimao Zhang , Sheng Wang , Zhen Li , Shuguang Cui

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings, that can be…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Taiki Miyanishi , Fumiya Kitamori , Shuhei Kurita , Jungdae Lee , Motoaki Kawanabe , Nakamasa Inoue

Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) methods for autonomous vehicles (AVs) typically struggle with ambiguous, context-dependent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haicheng Liao , Huanming Shen , Bonan Wang , Yongkang Li , Yihong Tang , Chengyue Wang , Dingyi Zhuang , Kehua Chen , Hai Yang , Chengzhong Xu , Zhenning Li

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Wenbin Tan , Jiawen Lin , Fangyong Wang , Yuan Xie , Yong Xie , Yachao Zhang , Yanyun Qu

Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. However, existing 3D understanding is predominantly based on 2D Vision-Language Models (VLMs),…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Runwei Guan , Jianan Liu , Ningwei Ouyang , Shaofeng Liang , Daizong Liu , Xiaolou Sun , Lianqing Zheng , Ming Xu , Yutao Yue , Guoqiang Mao , Hui Xiong

3D visual grounding aims to automatically locate the 3D region of the specified object given the corresponding textual description. Existing works fail to distinguish similar objects especially when multiple referred objects are involved in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Feng Xiao , Hongbin Xu , Qiuxia Wu , Wenxiong Kang

3D visual grounding consists of identifying the instance in a 3D scene which is referred by an accompanying language description. While several architectures have been proposed within the commonly employed grounding-by-selection framework,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Sombit Dey , Ozan Unal , Christos Sakaridis , Luc Van Gool

We introduce the task of localizing a flexible number of objects in real-world 3D scenes using natural language descriptions. Existing 3D visual grounding tasks focus on localizing a unique object given a text description. However, such a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yiming Zhang , ZeMing Gong , Angel X. Chang

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Yu Qi , Lipeng Gu , Honghua Chen , Liangliang Nan , Mingqiang Wei
‹ Prev 1 2 3 10 Next ›