中文
相关论文

相关论文: CityRefer: Geography-aware 3D Visual Grounding Dat…

200 篇论文

3D urban reconstruction of buildings from remotely sensed imagery has drawn significant attention during the past two decades. While aerial imagery and LiDAR provide higher resolution, satellite imagery is cheaper and more efficient to…

计算机视觉与模式识别 · 计算机科学 2020-05-20 Bo Xu , Xu Zhang , Zhixin Li , Matt Leotta , Shih-Fu Chang , Jie Shan

Large-scale terrain scans are the basis for many important tasks, such as topographic mapping, forestry, agriculture, and infrastructure planning. The resulting point cloud data sets are so massive in size that even basic tasks like viewing…

图形学 · 计算机科学 2025-09-25 Philipp Erler , Lukas Herzberger , Michael Wimmer , Markus Schütz

Vehicle localization using roadside LiDARs can provide centimeter-level accuracy for cloud-controlled vehicles while simultaneously serving multiple vehicles, enhanc-ing safety and efficiency. While most existing studies rely on repetitive…

机器人学 · 计算机科学 2025-09-22 Runxin Zhao , Chunxiang Wang , Hanyang Zhuang , Ming Yang

The construction industry increasingly relies on visual data to support Artificial Intelligence (AI) and Machine Learning (ML) applications for site monitoring. High-quality, domain-specific datasets, comprising images, videos, and point…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Ruoxin Xiong , Yanyu Wang , Jiannan Cai , Kaijian Liu , Yuansheng Zhu , Pingbo Tang , Nora El-Gohary

Access to labeled reference data is one of the grand challenges in supervised machine learning endeavors. This is especially true for an automated analysis of remote sensing images on a global scale, which enables us to address global…

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

The capability for open vocabulary perception represents a significant advancement in autonomous driving systems, facilitating the comprehension and interpretation of a wide array of textual inputs in real-time. Despite extensive research…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Xinlong Cheng , Lei Li

Semantic annotations are vital for training models for object recognition, semantic segmentation or scene understanding. Unfortunately, pixelwise annotation of images at very large scale is labor-intensive and only little labeled data is…

计算机视觉与模式识别 · 计算机科学 2016-04-13 Jun Xie , Martin Kiefel , Ming-Ting Sun , Andreas Geiger

The value of roadside perception, which could extend the boundaries of autonomous driving and traffic management, has gradually become more prominent and acknowledged in recent years. However, existing roadside perception approaches only…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Ruiyang Hao , Siqi Fan , Yingru Dai , Zhenlin Zhang , Chenxi Li , Yuntian Wang , Haibao Yu , Wenxian Yang , Jirui Yuan , Zaiqing Nie

Current point cloud registration methods are mainly based on local geometric information and usually ignore the semantic information contained in the scenes. In this paper, we treat the point cloud registration problem as a semantic…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Shaocong Liu , Tao Wang , Yan Zhang , Ruqin Zhou , Li Li , Chenguang Dai , Yongsheng Zhang , Longguang Wang , Hanyun Wang

Geolocation, the task of identifying an image's location, requires complex reasoning and is crucial for navigation, monitoring, and cultural preservation. However, current methods often produce coarse, imprecise, and non-interpretable…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Zirui Song , Jingpu Yang , Yuan Huang , Jonathan Tonglet , Zeyu Zhang , Tao Cheng , Meng Fang , Iryna Gurevych , Xiuying Chen

LiDAR perception is fundamental to robotics, enabling machines to understand their environment in 3D. A crucial task for LiDAR-based scene understanding and navigation is ground segmentation. However, existing methods are either handcrafted…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Ted Lentsch , Santiago Montiel-Marín , Holger Caesar , Dariu M. Gavrila

LiDAR is an important method for autonomous driving systems to sense the environment. The point clouds obtained by LiDAR typically exhibit sparse and irregular distribution, thus posing great challenges to the detection of 3D objects,…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Tai Wang , Xinge Zhu , Dahua Lin

LiDAR-produced point clouds are the major source for most state-of-the-art 3D object detectors. Yet, small, distant, and incomplete objects with sparse or few points are often hard to detect. We present Sparse2Dense, a new framework to…

计算机视觉与模式识别 · 计算机科学 2022-11-24 Tianyu Wang , Xiaowei Hu , Zhengzhe Liu , Chi-Wing Fu

3D point cloud segmentation aims to assign semantic labels to individual points in a scene for fine-grained spatial understanding. Existing methods typically adopt data augmentation to alleviate the burden of large-scale annotation.…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Hongbin Lin , Yifan Jiang , Juangui Xu , Jesse Jiaxi Xu , Yi Lu , Zhengyu Hu , Ying-Cong Chen , Hao Wang

Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on…

机器人学 · 计算机科学 2025-09-10 Yandi Yang , Jianping Li , Youqi Liao , Yuhao Li , Yizhe Zhang , Zhen Dong , Bisheng Yang , Naser El-Sheimy

Urban embodied AI agents, ranging from delivery robots to quadrupeds, are increasingly populating our cities, navigating chaotic streets to provide last-mile connectivity. Training such agents requires diverse, high-fidelity urban…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Mingxuan Liu , Honglin He , Elisa Ricci , Wayne Wu , Bolei Zhou

In this paper, we address semantic segmentation of road-objects from 3D LiDAR point clouds. In particular, we wish to detect and categorize instances of interest, such as cars, pedestrians and cyclists. We formulate this problem as a point-…

计算机视觉与模式识别 · 计算机科学 2017-10-23 Bichen Wu , Alvin Wan , Xiangyu Yue , Kurt Keutzer

Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scene-language models often struggle with this relational understanding,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jintang Xue , Ganning Zhao , Jie-En Yao , Hong-En Chen , Yue Hu , Meida Chen , Suya You , C. -C. Jay Kuo

The development of large-scale 3D scene reconstruction and novel view synthesis methods mostly rely on datasets comprising perspective images with narrow fields of view (FoV). While effective for small-scale scenes, these datasets require…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Ulas Gunes , Matias Turkulainen , Xuqian Ren , Arno Solin , Juho Kannala , Esa Rahtu