English
Related papers

Related papers: Global Cross-Modal Geo-Localization: A Million-Sca…

200 papers

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models (LVLMs) has enabled…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Ling Li , Yao Zhou , Yuxuan Liang , Fugee Tsung , Jiaheng Wei

The goal of point cloud localization based on linguistic description is to identify a 3D position using textual description in large urban environments, which has potential applications in various fields, such as determining the location…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Yanlong Xu , Haoxuan Qu , Jun Liu , Wenxiao Zhang , Xun Yang

Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zhiyang Lu , Wen Jiang , Tianren Wu , Zhichao Wang , Changwang Zhang , Siqi Shen , Ming Cheng

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual semantics, and 2) the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Pengyue Jia , Yiding Liu , Xiaopeng Li , Yuhao Wang , Yantong Du , Xiao Han , Xuetao Wei , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Place recognition is a critical component of autonomous vehicles and robotics, enabling global localization in GPS-denied environments. Recent advances have spurred significant interest in multimodal place recognition (MPR), which leverages…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Zhangshuo Qi , Jingyi Xu , Luqi Cheng , Shichen Wen , Yiming Ma , Guangming Xiong

Cross-modal retrieval has become a highlighted research topic for retrieval across multimedia data such as image and text. A two-stage learning framework is widely adopted by most existing methods based on Deep Neural Network (DNN): The…

Multimedia · Computer Science 2017-08-09 Yuxin Peng , Jinwei Qi , Xin Huang , Yuxin Yuan

Graph-based multi-view clustering aiming to obtain a partition of data across multiple views, has received considerable attention in recent years. Although great efforts have been made for graph-based multi-view clustering, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Yiming Wang , Dongxia Chang , Zhiqiang Fu , Yao Zhao

Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Xiaohan Zhang , Xingyu Li , Waqas Sultani , Yi Zhou , Safwan Wshah

Robust cross-view geo-localization (CVGL) remains challenging despite the surge in recent progress. Existing methods still rely on field-of-view (FoV)-specific training paradigms, where models are optimized under a fixed FoV but collapse…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yang Chen , Xieyuanli Chen , Junxiang Li , Jie Tang , Tao Wu

Existing LGL methods typically consider only partial information (e.g., geometric features) from LiDAR observations or are designed for homogeneous LiDAR sensors, overlooking the uniformity in LGL. In this work, a uniform LGL method is…

Robotics · Computer Science 2026-04-01 Hongming Shen , Xun Chen , Yulin Hui , Zhenyu Wu , Wei Wang , Qiyang Lyu , Tianchen Deng , Danwei Wang

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide geolocalization has been tackled before, but only using the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Parth Parag Kulkarni , Gaurav Kumar Nayak , Mubarak Shah

Reliable global localization is critical for autonomous vehicles, especially in environments where GNSS is degraded or unavailable, such as urban canyons and tunnels. Although high-definition (HD) maps provide accurate priors, the cost of…

Robotics · Computer Science 2025-09-30 Nguyen Hoang Khoi Tran , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Localization is a critically essential and crucial enabler of autonomous robots. While deep learning has made significant strides in many computer vision tasks, it is still yet to make a sizeable impact on improving capabilities of metric…

Computer Vision and Pattern Recognition · Computer Science 2020-05-25 Daniele Cattaneo , Domenico Giorgio Sorrenti , Abhinav Valada

Place recognition is one of the most crucial modules for autonomous vehicles to identify places that were previously visited in GPS-invalid environments. Sensor fusion is considered an effective method to overcome the weaknesses of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Zijie Zhou , Jingyi Xu , Guangming Xiong , Junyi Ma

Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address…

Machine Learning · Computer Science 2025-04-01 Jing Zhu , Yuhang Zhou , Shengyi Qian , Zhongmou He , Tong Zhao , Neil Shah , Danai Koutra

Omnidirectional scene understanding is vital for various downstream applications, such as embodied AI, autonomous driving, and immersive environments, yet remains challenging due to geometric distortion and complex spatial relations in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Xinshen Zhang , Tongxi Fu , Xu Zheng

Large Language Models (LLMs) are advanced deep-learning models designed to understand and generate human language. They work together with models that process data like images, enabling cross-modal understanding. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shreya Singh

Cross-view Geo-localisation is typically performed at a coarse granularity, because densely sampled satellite image patches overlap heavily. This heavy overlap would make disambiguating patches very challenging. However, by opting for…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Tavis Shore , Oscar Mendez , Simon Hadfield

Robots are increasingly operating in open-world environments where safe behavior depends on context: the same hallway may require different navigation strategies when crowded versus empty, or during an emergency versus normal operations.…