English
Related papers

Related papers: VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Pla…

200 papers

Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT) directly applies a pure transformer architecture on image…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Xiaoyu Yue , Shuyang Sun , Zhanghui Kuang , Meng Wei , Philip Torr , Wayne Zhang , Dahua Lin

Cross-view geo-localization (CVGL) aims to match images of the same location captured from drastically different viewpoints. Despite recent progress, existing methods still face two key challenges: (1) achieving robustness under severe…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Xiaowei Wang , Di Wang , Ke Li , Yifeng Wang , Chengjian Wang , Libin Sun , Zhihong Wu , Yiming Zhang , Quan Wang

We propose a novel end-to-end method for cross-view pose estimation. Given a ground-level query image and an aerial image that covers the query's local neighborhood, the 3 Degrees-of-Freedom camera pose of the query is estimated by matching…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Zimin Xia , Olaf Booij , Julian F. P. Kooij

Accurate localization is essential for autonomous driving, but GNSS-based methods struggle in challenging environments such as urban canyons. Cross-view pose optimization offers an effective solution by directly estimating vehicle pose…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Wooju Lee , Juhye Park , Dasol Hong , Changki Sung , Youngwoo Seo , Dongwan Kang , Hyun Myung

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Mengjie Zhou , Liu Liu , Yiran Zhong , Andrew Calway

This paper presents a visual geo-localization system capable of determining the geographic locations of places (buildings and road intersections) from images without relying on GPS data. Our approach integrates three primary methods:…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Rania Saoud , Slimane Larabi

Robust humanoid locomotion requires accurate and globally consistent perception of the surrounding 3D environment. However, existing perception modules, mainly based on depth images or elevation maps, offer only partial and locally…

Robotics · Computer Science 2025-11-19 Qingwei Ben , Botian Xu , Kailin Li , Feiyu Jia , Wentao Zhang , Jingping Wang , Jingbo Wang , Dahua Lin , Jiangmiao Pang

Accurate 3D perception is essential for understanding the environment in autonomous driving. Recent advancements in 3D semantic occupancy prediction have leveraged camera-LiDAR fusion to improve robustness and accuracy. However, current…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minjae Seong , Jisong Kim , Geonho Bang , Hawook Jeong , Jun Won Choi

Cross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Qiyuan Shen , Hengwang Zhao , Weihao Yan , Chunxiang Wang , Tong Qin , Ming Yang

Relative localization is critical for cooperation in autonomous multi-robot systems. Existing approaches either rely on shared environmental features or inertial assumptions or suffer from non-line-of-sight degradation and outliers in…

Robotics · Computer Science 2026-01-01 Zhehan Li , Zheng Wang , Jiadong Lu , Qi Liu , Zhiren Xun , Yue Wang , Fei Gao , Chao Xu , Yanjun Cao

Cross-model retrieval has emerged as one of the most important upgrades for text-only search engines (SE). Recently, with powerful representation for pairwise text-image inputs via early interaction, the accuracy of vision-language (VL)…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Lisai Zhang , Hongfa Wu , Qingcai Chen , Yimeng Deng , Zhonghua Li , Dejiang Kong , Zhao Cao , Joanna Siebert , Yunpeng Han

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Sifeng He , Yue He , Minlong Lu , Chen Jiang , Xudong Yang , Feng Qian , Xiaobo Zhang , Lei Yang , Jiandong Zhang

Visual Place Recognition (VPR) is a fundamental yet challenging task for small Unmanned Aerial Vehicle (UAV). The core reasons are the extreme viewpoint changes, and limited computational power onboard a UAV which restricts the…

Computer Vision and Pattern Recognition · Computer Science 2019-08-02 Bruno Ferrarini , Maria Waheed , Sania Waheed , Shoaib Ehsan , Michael Milford , Klaus D. McDonald-Maier

Vision-based bird's-eye-view (BEV) 3D object detection has advanced significantly in autonomous driving by offering cost-effectiveness and rich contextual information. However, existing methods often construct BEV representations by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jicheng Yuan , Manh Nguyen Duc , Qian Liu , Manfred Hauswirth , Danh Le Phuoc

Visual localization and mapping is a crucial capability to address many challenges in mobile robotics. It constitutes a robust, accurate and cost-effective approach for local and global pose estimation within prior maps. Yet, in highly…

Computer Vision and Pattern Recognition · Computer Science 2018-07-11 Guoxiang Zhou , Berta Bescos , Marcin Dymczyk , Mark Pfeiffer , José Neira , Roland Siegwart

Point clouds captured by different sensors such as RGB-D cameras and LiDAR possess non-negligible domain gaps. Most existing methods design different network architectures and train separately on point clouds from various sensors.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Shengjun Zhang , Xin Fei , Yueqi Duan

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensors. However,…

Robotics · Computer Science 2025-06-23 Junwei You , Haotian Shi , Zhuoyu Jiang , Zilin Huang , Rui Gan , Keshu Wu , Xi Cheng , Xiaopeng Li , Bin Ran

Roadside vision centric 3D object detection has received increasing attention in recent years. It expands the perception range of autonomous vehicles, enhances the road safety. Previous methods focused on predicting per-pixel height rather…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Zhang Zhang , Chao Sun , Chao Yue , Da Wen , Yujie Chen , Tianze Wang , Jianghao Leng

We present a novel method for local image feature matching. Instead of performing image feature detection, description, and matching sequentially, we propose to first establish pixel-wise dense matches at a coarse level and later refine the…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Jiaming Sun , Zehong Shen , Yuang Wang , Hujun Bao , Xiaowei Zhou

Recent Transformer-based 3D object detectors learn point cloud features either from point- or voxel-based representations. However, the former requires time-consuming sampling while the latter introduces quantization errors. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Honghui Yang , Wenxiao Wang , Minghao Chen , Binbin Lin , Tong He , Hua Chen , Xiaofei He , Wanli Ouyang