English
Related papers

Related papers: CityRefer: Geography-aware 3D Visual Grounding Dat…

200 papers

Place Recognition is a crucial capability for mobile robot localization and navigation. Image-based or Visual Place Recognition (VPR) is a challenging problem as scene appearance and camera viewpoint can change significantly when places are…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Sourav Garg , Michael Milford

As digital twins become central to the transformation of modern cities, accurate and structured 3D building models emerge as a key enabler of high-fidelity, updatable urban representations. These models underpin diverse applications…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Shangfeng Huang , Ruisheng Wang , Xin Wang

Urban waste management remains a critical challenge for the development of smart cities. Despite the growing number of litter detection datasets, the problem of monitoring overflowing waste containers, particularly from images captured by…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Diogo J. Paulo , João Martins , Hugo Proença , João C. Neves

3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence for describing each located object. However, the existing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Yufeng Zhong , Long Xu , Jiebo Luo , Lin Ma

Navigating large-scale outdoor environments requires complex reasoning in terms of geometric structures, environmental semantics, and terrain characteristics, which are typically captured by onboard sensors such as LiDAR and cameras. While…

Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predominantly focus on identifying and localizing objects from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhuofan Zhang , Ziyu Zhu , Junhao Li , Pengxiang Li , Tianxu Wang , Tengyu Liu , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Siyuan Huang , Qing Li

3D referring segmentation is an emerging and challenging vision-language task that aims to segment the object described by a natural language expression in a point cloud scene. The key challenge behind this task is vision-language feature…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Shuting He , Henghui Ding

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

Current state-of-the-art 3D reconstruction models face limitations in building extra-large scale outdoor scenes, primarily due to the lack of sufficiently large-scale and detailed datasets. In this paper, we present a extra-large…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xinyi Zheng , Steve Zhang , Weizhe Lin , Aaron Zhang , Walterio W. Mayol-Cuevas , Yunze Liu , Junxiao Shen

With the rapid advancement of 3D sensing technologies, obtaining 3D shape information of objects has become increasingly convenient. Lidar technology, with its capability to accurately capture the 3D information of objects at long…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Weixiao Gao , Ravi Peters , Jantien Stoter

To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded objects. However, current VG works are limited in working…

Computation and Language · Computer Science 2021-04-15 Ke-Jyun Wang , Yun-Hsuan Liu , Hung-Ting Su , Jen-Wei Wang , Yu-Siang Wang , Winston H. Hsu , Wen-Chin Chen

Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve…

Machine Learning · Computer Science 2019-06-04 Zhengping Che , Guangyu Li , Tracy Li , Bo Jiang , Xuefeng Shi , Xinsheng Zhang , Ying Lu , Guobin Wu , Yan Liu , Jieping Ye

Road scene understanding is crucial in autonomous driving, enabling machines to perceive the visual environment. However, recent object detectors tailored for learning on datasets collected from certain geographical locations struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Hasib Zunair , Shakib Khan , A. Ben Hamza

We present a fully automatic approach for reconstructing compact 3D building models from large-scale airborne point clouds. A major challenge of urban reconstruction from airborne LiDAR point clouds lies in that the vertical walls are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Jin Huang , Jantien Stoter , Ravi Peters , Liangliang Nan

This article presents UrbanTwin datasets, high-fidelity, realistic replicas of three public roadside lidar datasets: LUMPI, V2X-Real-IC, and TUMTraf-I. Each UrbanTwin dataset contains 10K annotated frames corresponding to one of the public…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Muhammad Shahbaz , Shaurya Agarwal

Due to the difficulty in generating the effective descriptors which are robust to occlusion and viewpoint changes, place recognition for 3D point cloud remains an open issue. Unlike most of the existing methods that focus on extracting…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Xin Kong , Xuemeng Yang , Guangyao Zhai , Xiangrui Zhao , Xianfang Zeng , Mengmeng Wang , Yong Liu , Wanlong Li , Feng Wen

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Henry Zheng , Hao Shi , Qihang Peng , Yong Xien Chng , Rui Huang , Yepeng Weng , Zhongchao Shi , Gao Huang

Street-level geolocalization from images is crucial for a wide range of essential applications and services, such as navigation, location-based recommendations, and urban planning. With the growing popularity of social media data and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yunus Serhat Bicakci , Joseph Shingleton , Anahid Basiri

Semantic understanding of the surrounding environment is essential for automated vehicles. The recent publication of the SemanticKITTI dataset stimulates the research on semantic segmentation of LiDAR point clouds in urban scenarios. While…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Juncong Fei , Kunyu Peng , Philipp Heidenreich , Frank Bieder , Christoph Stiller

Recently, there has been growing interest in developing learning-based methods to detect and utilize salient semi-global or global structures, such as junctions, lines, planes, cuboids, smooth surfaces, and all types of symmetries, for 3D…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Jia Zheng , Junfei Zhang , Jing Li , Rui Tang , Shenghua Gao , Zihan Zhou
‹ Prev 1 8 9 10 Next ›