English
Related papers

Related papers: CityRefer: Geography-aware 3D Visual Grounding Dat…

200 papers

The two popular datasets ScanRefer [16] and ReferIt3D [3] connect natural language to real-world 3D data. In this paper, we curate a large-scale and complementary dataset extending both the aforementioned ones by associating all objects…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Ahmed Abdelreheem , Kyle Olszewski , Hsin-Ying Lee , Peter Wonka , Panos Achlioptas

Semantic 3D city models are worldwide easy-accessible, providing accurate, object-oriented, and semantic-rich 3D priors. To date, their potential to mitigate the noise impact on radar object detection remains under-explored. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Yuan Luo , Rudolf Hoffmann , Yan Xia , Olaf Wysocki , Benedikt Schwab , Thomas H. Kolbe , Daniel Cremers

3D visual grounding is the ability to localize objects in 3D scenes conditioned by utterances. Most existing methods devote the referring head to localize the referred object directly, causing failure in complex scenarios. In addition, it…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Eslam Abdelrahman , Mohamed Ayman , Mahmoud Ahmed , Habib Slim , Mohamed Elhoseiny

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-04-25 Heng Wang , Chaoyi Zhang , Jianhui Yu , Weidong Cai

Advanced Driver-Assistance Systems (ADAS) have successfully integrated learning-based techniques into vehicle perception and decision-making. However, their application in 3D lane detection for effective driving environment perception is…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Runkai Zhao , Yuwen Heng , Heng Wang , Yuanda Gao , Shilei Liu , Changhao Yao , Jiawen Chen , Weidong Cai

Accurate, up-to-date High-Definition (HD) maps are critical for urban planning, infrastructure monitoring, and autonomous navigation. However, these maps quickly become outdated as environments evolve, creating a need for robust methods…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Chun-Jung Lin , Tat-Jun Chin , Sourav Garg , Feras Dayoub

3D semantic segmentation plays a critical role in urban modelling, enabling detailed understanding and mapping of city environments. In this paper, we introduce Turin3D: a new aerial LiDAR dataset for point cloud semantic segmentation…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Luca Barco , Giacomo Blanco , Gaetano Chiriaco , Alessia Intini , Luigi La Riccia , Vittorio Scolamiero , Piero Boccardo , Paolo Garza , Fabrizio Dominici

For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other. Recently, however, the community has realized that…

Computer Vision and Pattern Recognition · Computer Science 2022-06-06 Yiyi Liao , Jun Xie , Andreas Geiger

3D multi-object detection and tracking are crucial for traffic scene understanding. However, the community pays less attention to these areas due to the lack of a standardized benchmark dataset to advance the field. Moreover, existing…

Computer Vision and Pattern Recognition · Computer Science 2019-03-07 Abhishek Patil , Srikanth Malla , Haiming Gang , Yi-Ting Chen

Thanks to its precise spatial referencing, 3D point cloud visual grounding is essential for deep understanding and dynamic interaction in 3D environments, encompassing 3D Referring Expression Comprehension (3DREC) and Segmentation (3DRES).…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Haojia Lin , Yongdong Luo , Xiawu Zheng , Lijiang Li , Fei Chao , Taisong Jin , Donghao Luo , Yan Wang , Liujuan Cao , Rongrong Ji

3D visual grounding is an emerging research area dedicated to making connections between the 3D physical world and natural language, which is crucial for achieving embodied intelligence. In this paper, we propose DASANet, a Dual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Yue Xu , Kaizhi Yang , Jiebo Luo , Xuejin Chen

Point cloud-based open-vocabulary 3D object detection aims to detect 3D categories that do not have ground-truth annotations in the training set. It is extremely challenging because of the limited data and annotations (bounding boxes with…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chenming Zhu , Wenwei Zhang , Tai Wang , Xihui Liu , Kai Chen

Traditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 HyunJun Jung , Weihang Li , Shun-Cheng Wu , William Bittner , Nikolas Brasch , Jifei Song , Eduardo Pérez-Pellitero , Zhensong Zhang , Arthur Moreau , Nassir Navab , Benjamin Busam

Determining the location of an image anywhere on Earth is a complex visual task, which makes it particularly relevant for evaluating computer vision algorithms. Yet, the absence of standard, large-scale, open-access datasets with reliably…

Identifying and classifying underground utilities is an important task for efficient and effective urban planning and infrastructure maintenance. We present OpenTrench3D, a novel and comprehensive 3D Semantic Segmentation point cloud…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Lasse H. Hansen , Simon B. Jensen , Mark P. Philipsen , Andreas Møgelmose , Lars Bodum , Thomas B. Moeslund

Large collections of geo-referenced panoramic images are freely available for cities across the globe, as well as detailed maps with location and meta-data on a great variety of urban objects. They provide a potentially rich source of…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Inske Groenen , Stevan Rudinac , Marcel Worring

3D visual grounding aims to automatically locate the 3D region of the specified object given the corresponding textual description. Existing works fail to distinguish similar objects especially when multiple referred objects are involved in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Feng Xiao , Hongbin Xu , Qiuxia Wu , Wenxiong Kang

Research on supervised learning algorithms in 3D scene understanding has risen in prominence and witness great increases in performance across several datasets. The leading force of this research is the problem of autonomous driving…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Diogo Lavado , Cláudia Soares , Alessandra Micheletti , Ricardo Santos , André Coelho , João Santos

Visual grounding aims to identify objects or regions in a scene based on natural language descriptions, essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically depend on bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Zhan Shi , Song Wang , Junbo Chen , Jianke Zhu

Detecting 3D objects keypoints is of great interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either…

Computer Vision and Pattern Recognition · Computer Science 2020-08-10 Yang You , Yujing Lou , Chengkun Li , Zhoujun Cheng , Liangwei Li , Lizhuang Ma , Weiming Wang , Cewu Lu