English
Related papers

Related papers: ConGeo: Robust Cross-view Geo-localization across …

200 papers

Today's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human learning. Can we make…

Computation and Language · Computer Science 2024-03-22 Chengxu Zhuang , Evelina Fedorenko , Jacob Andreas

Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigation and emergency response. However, existing researches are constrained by narrow geographic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yutong Hu , Jinhui Chen , Chaoqiang Xu , Yuan Kou , Sili Zhou , Shaocheng Yan , Pengcheng Shi , Qingwu Hu , Jiayuan Li

Feed-forward multi-frame 3D reconstruction models often degrade on videos with object motion. Global-reference becomes ambiguous under multiple motions, while the local pointmap relies heavily on estimated relative poses and can drift,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Xingyu Miao , Weiguang Zhao , Tao Lu , Linning Xu , Mulin Yu , Yang Long , Jiangmiao Pang , Junting Dong

Vision-language models (VLMs) enable text-guided object detection but degrade severely under cross-view scenarios where ground and aerial viewpoints differ in altitude, scale, and spatial layout. These geometric changes introduce systematic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhipeng Liu , Chunbo Luo

Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Yongfei Liu , Bo Wan , Lin Ma , Xuming He

The ability to evolve is fundamental for any valuable autonomous agent whose knowledge cannot remain limited to that injected by the manufacturer. Consider for example a home assistant robot: it should be able to incrementally learn new…

Computer Vision and Pattern Recognition · Computer Science 2022-09-05 Francesco Cappio Borlino , Silvia Bucci , Tatiana Tommasi

Cross-view geo-localization is a task of matching the same geographic image from different views, e.g., unmanned aerial vehicle (UAV) and satellite. The most difficult challenges are the position shift and the uncertainty of distance and…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Ming Dai , Jianhong Hu , Jiedong Zhuang , Enhui Zheng

Monocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a training-free Monocular Geometry Refinement method designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Dongki Jung , Jaehoon Choi , Yonghan Lee , Sungmin Eum , Heesung Kwon , Dinesh Manocha

We consider a class of inverse problems characterized by forward operators that are partially specified, non-smooth, and non-differentiable. Although generative inverse solvers have made significant progress, we find that these forward…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Sattwik Basu , Chaitanya Amballa , Zhongweiyang Xu , Jorge Vančo Sampedro , Srihari Nelakuditi , Romit Roy Choudhury

Visual Geo-localization (VG) is the task of estimating the position where a given photo was taken by comparing it with a large database of images of known locations. To investigate how existing techniques would perform on a real-world…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Gabriele Berton , Carlo Masone , Barbara Caputo

Video grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-08 Guoshun Nan , Rui Qiao , Yao Xiao , Jun Liu , Sicong Leng , Hao Zhang , Wei Lu

Object detection has advanced significantly in the closed-set setting, but real-world deployment remains limited by two challenges: poor generalization to unseen categories and insufficient robustness under adverse conditions. Prior…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Siheng Wang , Zhengdao Li , Yanshu Li , Canran Xiao , Haibo Zhan , Zhengtao Yao , Xuzhi Zhang , Jiale Kang , Linshan Li , Weiming Liu , Zhikang Dong , Jifeng Shen , Junhao Dong , Qiang Sun , Piotr Koniusz

Bird's-eye-view (BEV) representations derived from multi-camera input have become a central interface for online high-definition (HD) map construction. However, most approaches rely solely on ego-centric supervision, requiring large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Daniel Lengerer , Mathias Pechinger , Klaus Bogenberger , Carsten Markgraf

As Earth's climate changes, it is impacting disasters and extreme weather events across the planet. Record-breaking heat waves, drenching rainfalls, extreme wildfires, and widespread flooding during hurricanes are all becoming more frequent…

Artificial Intelligence · Computer Science 2026-05-07 Hao Li , Fabian Deuser , Wenping Yin , Steffen Knoblauch , Wufan Zhao , Filip Biljecki , Yong Xue , Wei Huang

In this paper, we address the problem of global-scale image geolocation, proposing a mixed classification-retrieval scheme. Unlike other methods that strictly tackle the problem as a classification or retrieval task, we combine the two…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Giorgos Kordopatis-Zilos , Panagiotis Galopoulos , Symeon Papadopoulos , Ioannis Kompatsiaris

In this paper, we introduce a novel approach to fine-grained cross-view geo-localization. Our method aligns a warped ground image with a corresponding GPS-tagged satellite image covering the same area using homography estimation. We first…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Xiaolong Wang , Runsen Xu , Zuofan Cui , Zeyu Wan , Yu Zhang

We propose to jointly learn multi-view geometry and warping between views of the same object instances for robust cross-view object detection. What makes multi-view object instance detection difficult are strong changes in viewpoint,…

Machine Learning · Computer Science 2019-07-26 Ahmed Samy Nassar , Sebastien Lefevre , Jan D. Wegner

The ground-to-satellite image matching/retrieval was initially proposed for city-scale ground camera localization. This work addresses the problem of improving camera pose accuracy by ground-to-satellite image matching after a coarse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Yujiao Shi , Hongdong Li , Akhil Perincherry , Ankit Vora

Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful technique than object detection in practice. Most…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Chaorui Deng , Qi Wu , Guanghui Xu , Zhuliang Yu , Yanwu Xu , Kui Jia , Mingkui Tan

Weakly supervised localization aims at finding target object regions using only image-level supervision. However, localization maps extracted from classification networks are often not accurate due to the lack of fine pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Xiaolin Zhang , Yunchao Wei , Yi Yang