中文
相关论文

相关论文: Global Cross-Modal Geo-Localization: A Million-Sca…

200 篇论文

Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Angel Daruna , Nicholas Meegan , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus (satellite/drone/ground) given a natural-language query. Two…

计算机视觉与模式识别 · 计算机科学 2025-10-24 LinFeng Li , Jian Zhao , Zepeng Yang , Yuhang Song , Bojun Lin , Tianle Zhang , Yuchen Yuan , Chi Zhang , Xuelong Li

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Multimodal Contrastive Learning (MCL) advances in aligning different modalities and generating multimodal representations in a joint space. By leveraging contrastive learning across diverse modalities, large-scale multimodal data enhances…

机器学习 · 计算机科学 2025-09-23 Xiaohao Liu , Xiaobo Xia , See-Kiong Ng , Tat-Seng Chua

Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on…

机器人学 · 计算机科学 2025-09-10 Yandi Yang , Jianping Li , Youqi Liao , Yuhao Li , Yizhe Zhang , Zhen Dong , Bisheng Yang , Naser El-Sheimy

Land use and land cover mapping from Earth Observation (EO) data is a critical tool for sustainable land and resource management. While advanced machine learning and deep learning algorithms excel at analyzing EO imagery data, they often…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Babak Ghassemi , Cassio Fraga-Dantas , Raffaele Gaetano , Dino Ienco , Omid Ghorbanzadeh , Emma Izquierdo-Verdiguier , Francesco Vuolo

Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and architectural styles. Despite advancements in geo-localization…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Furong Jia , Lanxin Liu , Ce Hou , Fan Zhang , Xinyan Liu , Yu Liu

Representation learning of geospatial locations remains a core challenge in achieving general geospatial intelligence, with increasingly diverging philosophies and techniques. While Earth observation paradigms excel at depicting locations…

人工智能 · 计算机科学 2026-01-28 Ya Wen , Jixuan Cai , Qiyao Ma , Linyan Li , Xinhua Chen , Chris Webster , Yulun Zhou

Existing deep learning-based cross-view geo-localization methods primarily focus on improving the accuracy of cross-domain image matching, rather than enabling models to comprehensively capture contextual information around the target and…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Suofei Zhang , Xinxin Wang , Xiaofu Wu , Quan Zhou , Haifeng Hu

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic…

计算机视觉与模式识别 · 计算机科学 2019-11-26 Yongfei Liu , Bo Wan , Xiaodan Zhu , Xuming He

We present a visual localization system that learns to estimate camera poses in the real world with the help of synthetic data. Despite significant progress in recent years, most learning-based approaches to visual localization target at a…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Qi Yan , Jianhao Zheng , Simon Reding , Shanci Li , Iordan Doytchinov

We present GLEE in this work, an object-level foundation model for locating and identifying objects in images and videos. Through a unified framework, GLEE accomplishes detection, segmentation, tracking, grounding, and identification of…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Junfeng Wu , Yi Jiang , Qihao Liu , Zehuan Yuan , Xiang Bai , Song Bai

Multimodal learning plays a pivotal role in advancing artificial intelligence systems by incorporating information from multiple modalities to build a more comprehensive representation. Despite its importance, current state-of-the-art…

机器学习 · 计算机科学 2025-09-30 Giordano Cicchetti , Eleonora Grassucci , Danilo Comminiello

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart,…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xuanke Shi , Boxuan Li , Xiaoyang Han , Zhongang Cai , Lei Yang , Quan Wang , Dahua Lin

Cross-modal retrieval aims to measure the content similarity between different types of data. The idea has been previously applied to visual, text, and speech data. In this paper, we present a novel cross-modal retrieval method specifically…

计算机视觉与模式识别 · 计算机科学 2020-05-05 Numan Khurshid , Talha Hanif , Mohbat Tharani , Murtaza Taj

Multi-relational graph clustering has demonstrated remarkable success in uncovering underlying patterns in complex networks. Representative methods manage to align different views motivated by advances in contrastive learning. Our empirical…

机器学习 · 计算机科学 2024-07-25 Zhixiang Shen , Haolan He , Zhao Kang

Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hallucination, attention leakage, and failures in object…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Lihao Zheng , Zhenwei Shao , Yu Zhou , Yan Yang , Xintian Shen , Jiawei Chen , Hao Ma , Tao Wei

Graphs are powerful representations for relations among objects, which have attracted plenty of attention. A fundamental challenge for graph learning is how to train an effective Graph Neural Network (GNN) encoder without labels, which are…

机器学习 · 计算机科学 2022-10-19 Baoyu Jing , Shengyu Feng , Yuejia Xiang , Xi Chen , Yu Chen , Hanghang Tong

Precise estimation of global orientation and location is critical to ensure a compelling outdoor Augmented Reality (AR) experience. We address the problem of geo-pose estimation by cross-view matching of query ground images to a…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Niluthpol Chowdhury Mithun , Kshitij Minhas , Han-Pang Chiu , Taragay Oskiper , Mikhail Sizintsev , Supun Samarasekera , Rakesh Kumar

To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allowing seamless switching between map and query sources. It…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yan Xia , Zhendong Li , Yun-Jin Li , Letian Shi , Hu Cao , João F. Henriques , Daniel Cremers