English
Related papers

Related papers: Global Cross-Modal Geo-Localization: A Million-Sca…

200 papers

Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spatial-structure discrepancies between ground observations and aerial references. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhengyi Xu , Yuhang Ming , Zhihao Zhan , Hanyu Zhu , Javier Civera , Wanzeng Kong

Robot localization remains a challenging task in GPS denied environments. State estimation approaches based on local sensors, e.g. cameras or IMUs, are drifting-prone for long-range missions as error accumulates. In this study, we aim to…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Tianyi Zhang , Matthew Johnson-Roberson

Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jingyu Lei , Gaoang Wang , Der-Horng Lee

Graph Foundation Models (GFMs) have achieved remarkable success in generalizing across diverse domains. However, they mainly focus on Text-Attributed Graphs (TAGs), leaving Multimodal-Attributed Graphs (MAGs) largely untapped. Developing…

Machine Learning · Computer Science 2026-02-05 Sicheng Liu , Xunkai Li , Daohan Su , Ru Zhang , Hongchao Qin , Ronghua Li , Guoren Wang

Recent advances in open-vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align language and vision modalities. However, these approaches often neglect internal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Bozhao Li , Shaocong Wu , Tong Shao , Senqiao Yang , Qiben Shan , Zhuotao Tian , Jingyong Su

Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques. We propose utilising graph representations to model sequences of local observations…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Tavis Shore , Oscar Mendez , Simon Hadfield

In autonomous driving, robust place recognition is critical for global localization and loop closure detection. While inter-modality fusion of camera and LiDAR data in multimodal place recognition (MPR) has shown promise in overcoming the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Jingyi Xu , Zhangshuo Qi , Zhongmiao Yan , Xuyu Gao , Qianyun Jiao , Songpengcheng Xia , Xieyuanli Chen , Ling Pei

Contextual information is vital in visual understanding problems, such as semantic segmentation and object detection. We propose a Criss-Cross Network (CCNet) for obtaining full-image contextual information in a very effective and efficient…

Computer Vision and Pattern Recognition · Computer Science 2020-07-10 Zilong Huang , Xinggang Wang , Yunchao Wei , Lichao Huang , Humphrey Shi , Wenyu Liu , Thomas S. Huang

Visual Geo-localization (VG) refers to the process to identify the location described in query images, which is widely applied in robotics field and computer vision tasks, such as autonomous driving, metaverse, augmented reality, and SLAM.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Chen Mao , Jingqi Hu

Determining the exact latitude and longitude that a photo was taken is a useful and widely applicable task, yet it remains exceptionally difficult despite the accelerated progress of other computer vision tasks. Most previous approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Brandon Clark , Alec Kerrigan , Parth Parag Kulkarni , Vicente Vivanco Cepeda , Mubarak Shah

We propose a multi-camera multi-target (MCMT) tracking framework that ensures consistent global identity assignment across views using trajectory and appearance cues. The pipeline starts with BoT-SORT-based single-camera tracking, followed…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Hamidreza Hashempoor

This paper presents Vision-Language Global Localization (VLG-Loc), a novel global localization method that uses human-readable labeled footprint maps containing only names and areas of distinctive visual landmarks in an environment. While…

Robotics · Computer Science 2025-12-19 Mizuho Aoki , Kohei Honda , Yasuhiro Yoshimura , Takeshi Ishita , Ryo Yonetani

Causal Representation Learning (CRL) aims to uncover the data-generating process and identify the underlying causal variables and relations, whose evaluation remains inherently challenging due to the requirement of known ground-truth causal…

Machine Learning · Computer Science 2025-10-20 Guangyi Chen , Yunlong Deng , Peiyuan Zhu , Yan Li , Yifan Shen , Zijian Li , Kun Zhang

Knowledge graph completion (KGC) aims to automatically infer missing facts in multi-relational data by mapping entities and relations into continuous representation spaces. Recent region-based embedding models have shown great promise in…

Machine Learning · Computer Science 2026-05-13 Yingqi Zeng , Luying Wang , Huiling Zhu

Recent years have seen a significant increase in demand for robotic solutions in unstructured natural environments, alongside growing interest in bridging 2D and 3D scene understanding. However, existing robotics datasets are predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Joshua Knights , Joseph Reid , Kaushik Roy , David Hall , Mark Cox , Peyman Moghadam

We consider the problem of cross-view geo-localization. The primary challenge of this task is to learn the robust feature against large viewpoint changes. Existing benchmarks can help, but are limited in the number of viewpoints. Image…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Zhedong Zheng , Yunchao Wei , Yi Yang

This paper presents the multi-modal BigEarthNet (BigEarthNet-MM) benchmark archive made up of 590,326 pairs of Sentinel-1 and Sentinel-2 image patches to support the deep learning (DL) studies in multi-modal multi-label remote sensing (RS)…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Gencer Sumbul , Arne de Wall , Tristan Kreuziger , Filipe Marcelino , Hugo Costa , Pedro Benevides , Mário Caetano , Begüm Demir , Volker Markl

Place recognition is a challenging task in computer vision, crucial for enabling autonomous vehicles and robots to navigate previously visited environments. While significant progress has been made in learnable multimodal methods that…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Alexander Melekhin , Dmitry Yudin , Ilia Petryashin , Vitaly Bezuglyj

Several studies rely on the de facto standard Adaptive Monte Carlo Localization (AMCL) method to localize a robot in an Occupancy Grid Map (OGM) extracted from a building information model (BIM model). However, most of these studies assume…

Robotics · Computer Science 2023-08-11 Miguel Arturo Vega Torres , Alexander Braun , André Borrmann

The primary goal of artificial intelligence is to mimic humans. Therefore, to advance toward this goal, the AI community attempts to imitate qualities/skills possessed by humans and imbibes them into machines with the help of…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Binoy Saha , Sukhendu Das