English
Related papers

Related papers: Cross3DVG: Cross-Dataset 3D Visual Grounding on Di…

200 papers

Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promising results but rely on view-dependent reasoning or implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xuefei Sun , Xujia Zhang , Brendan Crowe , Doncey Albin , Christoffer Heckman

Access to large, diverse RGB-D datasets is critical for training RGB-D scene understanding algorithms. However, existing datasets still cover only a limited number of views or a restricted scale of spaces. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2017-09-20 Angel Chang , Angela Dai , Thomas Funkhouser , Maciej Halber , Matthias Nießner , Manolis Savva , Shuran Song , Andy Zeng , Yinda Zhang

Visual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the target from other objects, particularly for the targets that…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Wei Tang , Liang Li , Xuejing Liu , Lu Jin , Jinhui Tang , Zechao Li

The availability of vast amounts of visual data with heterogeneous features is a key factor for developing, testing, and benchmarking of new computer vision (CV) algorithms and architectures. Most visual datasets are created and curated for…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Jicheng Yuan , Anh Le-Tuan , Manh Nguyen-Duc , Trung-Kien Tran , Manfred Hauswirth , Danh Le-Phuoc

3D Visual Grounding (3DVG) aims to localize target objects within a 3D scene based on natural language queries. To alleviate the reliance on costly 3D training data, recent studies have explored zero-shot 3DVG by leveraging the extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Zhao Jin , Rong-Cheng Tu , Jingyi Liao , Wenhao Sun , Xiao Luo , Shunyu Liu , Dacheng Tao

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Zehan Wang , Haifeng Huang , Yang Zhao , Linjun Li , Xize Cheng , Yichen Zhu , Aoxiong Yin , Zhou Zhao

We introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing, the direct capture of the depth channel allows better 3D annotations…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Hongchi Xia , Yang Fu , Sifei Liu , Xiaolong Wang

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to understand video. To…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Hong Zhou , Mike Zheng Shou , Xiang Bai

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

Most advanced visual grounding methods rely on Transformers for visual-linguistic feature fusion. However, these Transformer-based approaches encounter a significant drawback: the computational costs escalate quadratically due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Wei Chen , Long Chen , Yu Wu

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to capturing the 3D visual…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Eslam Mohamed Bakr , Yasmeen Alsaedy , Mohamed Elhoseiny

Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radiology reporting. However, these models require large…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Zachary Huemann , Samuel Church , Joshua D. Warner , Daniel Tran , Xin Tie , Alan B McMillan , Junjie Hu , Steve Y. Cho , Meghan Lubner , Tyler J. Bradshaw

Visual Grounding (VG) aims at localizing target objects from an image based on given expressions and has made significant progress with the development of detection and vision transformer. However, existing VG methods tend to generate…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Menghao Li , Chunlei Wang , Wenquan Feng , Shuchang Lyu , Guangliang Cheng , Xiangtai Li , Binghao Liu , Qi Zhao

Visual grounding aims to predict the locations of target objects specified by textual descriptions. For this task with linguistic and visual modalities, there is a latest research line that focuses on only selecting the linguistic-relevant…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jingchao Wang , Wenlong Zhang , Dingjiang Huang , Hong Wang , Yefeng Zheng

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Haochen Wang , Yucheng Zhao , Tiancai Wang , Haoqiang Fan , Xiangyu Zhang , Zhaoxiang Zhang

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xueying Jiang , Lewei Lu , Ling Shao , Shijian Lu

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Ming Dai , Wenxuan Cheng , Jiang-Jiang Liu , Lingfeng Yang , Zhenhua Feng , Wankou Yang , Jingdong Wang

3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Due to its low cost and high efficiency, multi-view 3D object detection has demonstrated promising application prospects.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Zehui Chen , Zhenyu Li , Shiquan Zhang , Liangji Fang , Qinhong Jiang , Feng Zhao

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Haochen Wang , Xiangtai Li , Zilong Huang , Anran Wang , Jiacong Wang , Tao Zhang , Jiani Zheng , Sule Bai , Zijian Kang , Jiashi Feng , Zhuochen Wang , Zhaoxiang Zhang

The advancement of 3D vision-language (3D VL) learning is hindered by several limitations in existing 3D VL datasets: they rarely necessitate reasoning beyond a close range of objects in single viewpoint, and annotations often link…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Wentao Mo , Qingchao Chen , Yuxin Peng , Siyuan Huang , Yang Liu