English
Related papers

Related papers: CityRefer: Geography-aware 3D Visual Grounding Dat…

200 papers

Precise segmentation of architectural structures provides detailed information about various building components, enhancing our understanding and interaction with our built environment. Nevertheless, existing outdoor 3D point cloud datasets…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Ka Lung Cheung , Chi Chung Lee

Understanding the complex urban infrastructure with centimeter-level accuracy is essential for many applications from autonomous driving to mapping, infrastructure monitoring, and urban management. Aerial images provide valuable information…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Seyed Majid Azimi , Corentin Henry , Lars Sommer , Arne Schumann , Eleonora Vig

This paper addresses the problem of 3D referring expression comprehension (REC) in autonomous driving scenario, which aims to ground a natural language to the targeted region in LiDAR point clouds. Previous approaches for REC usually focus…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Wenhao Cheng , Junbo Yin , Wei Li , Ruigang Yang , Jianbing Shen

Current 3D visual grounding tasks only process sentence level detection or segmentation, which critically fails to leverage the rich compositional contextual reasonings within natural language expressions. To address this challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Qi Chen , Changli Wu , Jiayi Ji , Yiwei Ma , Liujuan Cao

Open-vocabulary semantic segmentation enables models to recognize and segment objects from arbitrary natural language descriptions, offering the flexibility to handle novel, fine-grained, or functionally defined categories beyond fixed…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Chongyu Wang , Kunlei Jing , Jihua Zhu , Di Wang

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Xinyi Wang , Na Zhao , Zhiyuan Han , Dan Guo , Xun Yang

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Haochen Zhang , Nader Zantout , Pujith Kachana , Ji Zhang , Wenshan Wang

We introduce a new outdoor urban 3D pointcloud dataset, covering a total area of 2.7 $km^2$, sampled from three Swiss cities with different characteristics. The dataset is manually annotated for semantic segmentation with per-point labels,…

Computer Vision and Pattern Recognition · Computer Science 2020-12-25 Gülcan Can , Dario Mantegazza , Gabriele Abbate , Sébastien Chappuis , Alessandro Giusti

Semantic segmentation of large-scale outdoor point clouds is essential for urban scene understanding in various applications, especially autonomous driving and urban high-definition (HD) mapping. With rapid developments of mobile laser…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Weikai Tan , Nannan Qin , Lingfei Ma , Ying Li , Jing Du , Guorong Cai , Ke Yang , Jonathan Li

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Mingtao Feng , Zhen Li , Qi Li , Liang Zhang , XiangDong Zhang , Guangming Zhu , Hui Zhang , Yaonan Wang , Ajmal Mian

We introduce BuildingNet: (a) a large-scale dataset of 3D building models whose exteriors are consistently labeled, (b) a graph neural network that labels building meshes by analyzing spatial and structural relations of their geometric…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Pratheba Selvaraju , Mohamed Nabail , Marios Loizou , Maria Maslioukova , Melinos Averkiou , Andreas Andreou , Siddhartha Chaudhuri , Evangelos Kalogerakis

Detecting pedestrians is a crucial task in autonomous driving systems to ensure the safety of drivers and pedestrians. The technologies involved in these algorithms must be precise and reliable, regardless of environment conditions. Relying…

Computer Vision and Pattern Recognition · Computer Science 2021-05-05 Òscar Lorente , Josep R. Casas , Santiago Royo , Ivan Caminal

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to capturing the 3D visual…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Eslam Mohamed Bakr , Yasmeen Alsaedy , Mohamed Elhoseiny

Recent progress in 3D scene understanding has explored visual grounding (3DVG) to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Zhihao Yuan , Xu Yan , Zhuo Li , Xuhao Li , Yao Guo , Shuguang Cui , Zhen Li

High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and asset monitoring, including both built structures and urban greenery. Conventional Digital Surface…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Hezam Albagami , Haitian Wang , Xinyu Wang , Muhammad Ibrahim , Zainy M. Malakan , Abdullah M. Alqamdi , Mohammed H. Alghamdi , Ajmal Mian

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide geolocalization has been tackled before, but only using the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Parth Parag Kulkarni , Gaurav Kumar Nayak , Mubarak Shah

Intelligent Transportation Systems (ITS) allow a drastic expansion of the visibility range and decrease occlusions for autonomous driving. To obtain accurate detections, detailed labeled sensor data for training is required. Unfortunately,…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Walter Zimmer , Christian Creß , Huu Tung Nguyen , Alois C. Knoll

With the enhancement of remote sensing image resolution and the rapid advancement of deep learning, land cover mapping is transitioning from pixel-level segmentation to object-based vector modeling. This shift demands more from deep…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Yu Meng , Ligao Deng , Zhihao Xi , Jiansheng Chen , Jingbo Chen , Anzhi Yue , Diyou Liu , Kai Li , Chenhao Wang , Kaiyu Li , Yupeng Deng , Xian Sun

Video-based gait recognition has achieved impressive results in constrained scenarios. However, visual cameras neglect human 3D structure information, which limits the feasibility of gait recognition in the 3D wild world. Instead of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Chuanfu Shen , Chao Fan , Wei Wu , Rui Wang , George Q. Huang , Shiqi Yu

Place recognition plays an essential role in the field of autonomous driving and robot navigation. Point cloud based methods mainly focus on extracting global descriptors from local features of point clouds. Despite having achieved…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Tian-Xing Xu , Yuan-Chen Guo , Zhiqiang Li , Ge Yu , Yu-Kun Lai , Song-Hai Zhang