English
Related papers

Related papers: CityRefer: Geography-aware 3D Visual Grounding Dat…

200 papers

Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. However, existing 3D understanding is predominantly based on 2D Vision-Language Models (VLMs),…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Runwei Guan , Jianan Liu , Ningwei Ouyang , Shaofeng Liang , Daizong Liu , Xiaolou Sun , Lianqing Zheng , Ming Xu , Yutao Yue , Guoqiang Mao , Hui Xiong

To meet the challenges of global urbanization, earth observation information is greatly needed. The lack of global three-dimensional (3D) urban structure data has been a major limiting factor in important urban applications such as…

Physics and Society · Physics 2018-07-13 Panshi Wang , Chengquan Huang , James C. Tilton

The development of computer vision algorithms for Unmanned Aerial Vehicles (UAVs) imagery heavily relies on the availability of annotated high-resolution aerial data. However, the scarcity of large-scale real datasets with pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Giulia Rizzoli , Francesco Barbato , Matteo Caligiuri , Pietro Zanuttigh

3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Baoxiong Jia , Yixin Chen , Huangyue Yu , Yan Wang , Xuesong Niu , Tengyu Liu , Qing Li , Siyuan Huang

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zirui Wang , Tao Zhang

Humans can orient themselves in their 3D environments using simple 2D maps. Differently, algorithms for visual localization mostly rely on complex 3D point clouds that are expensive to build, store, and maintain over time. We bridge this…

When performing localization and mapping, working at the level of structure can be advantageous in terms of robustness to environmental changes and differences in illumination. This paper presents SegMap: a map representation solution to…

Robotics · Computer Science 2019-01-16 Renaud Dubé , Andrei Cramariuc , Daniel Dugas , Juan Nieto , Roland Siegwart , Cesar Cadena

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuting He , Guangquan Jie , Changshuo Wang , Yun Zhou , Shuming Hu , Guanbin Li , Henghui Ding

Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a…

We are interested in understanding whether retrieval-based localization approaches are good enough in the context of self-driving vehicles. Towards this goal, we introduce Pit30M, a new image and LiDAR dataset with over 30 million frames,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-02 Julieta Martinez , Sasha Doubov , Jack Fan , Ioan Andrei Bârsan , Shenlong Wang , Gellért Máttyus , Raquel Urtasun

Semantic segmentation in urban scene analysis has mainly focused on images or point clouds, while textured meshes - offering richer spatial representation - remain underexplored. This paper introduces SUM Parts, the first large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Weixiao Gao , Liangliang Nan , Hugo Ledoux

Over the past few years, there has been remarkable progress in research on 3D point clouds and their use in autonomous driving scenarios has become widespread. However, deep learning methods heavily rely on annotated data and often face…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Jin Fang , Dingfu Zhou , Jingjing Zhao , Chenming Wu , Chulin Tang , Cheng-Zhong Xu , Liangjun Zhang

Recent developments in data acquisition technology allow us to collect 3D texture meshes quickly. Those can help us understand and analyse the urban environment, and as a consequence are useful for several applications like spatial analysis…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Weixiao Gao , Liangliang Nan , Bas Boom , Hugo Ledoux

With deep learning becoming a more prominent approach for automatic classification of three-dimensional point cloud data, a key bottleneck is the amount of high quality training data, especially when compared to that available for…

Computer Vision and Pattern Recognition · Computer Science 2019-07-11 David Griffiths , Jan Boehm

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress in 3D part understanding, their reliance on untextured…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Penghao Wang , Yiyang He , Xin Lv , Yukai Zhou , Lan Xu , Jingyi Yu , Jiayuan Gu

We present the UrbanBIS benchmark for large-scale 3D urban understanding, supporting practical urban-level semantic and building-level instance segmentation. UrbanBIS comprises six real urban scenes, with 2.5 billion points, covering a vast…

Graphics · Computer Science 2023-05-05 Guoqing Yang , Fuyou Xue , Qi Zhang , Ke Xie , Chi-Wing Fu , Hui Huang

Autonomous vehicles operate in highly dynamic environments necessitating an accurate assessment of which aspects of a scene are moving and where they are moving to. A popular approach to 3D motion estimation, termed scene flow, is to employ…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Philipp Jund , Chris Sweeney , Nichola Abdo , Zhifeng Chen , Jonathon Shlens

Reflective surfaces present a persistent challenge for reliable 3D mapping and perception in robotics and autonomous systems. However, existing reflection datasets and benchmarks remain limited to sparse 2D data. This paper introduces the…

Robotics · Computer Science 2024-03-12 Xiting Zhao , Sören Schwertfeger

3D semantic scene understanding remains a long-standing challenge in the 3D computer vision community. One of the key issues pertains to limited real-world annotated data to facilitate generalizable models. The common practice to tackle…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Duc Nguyen , Yan-Ling Lai , Qilin Zhang , Prabin Gyawali , Benedikt Schwab , Olaf Wysocki , Thomas H. Kolbe

This paper introduces DensePoint, a densely sampled and annotated point cloud dataset containing over 10,000 single objects across 16 categories, by merging different kind of information from two existing datasets. Each point cloud in…

Computer Vision and Pattern Recognition · Computer Science 2018-10-15 Xu Cao , Katashi Nagao