English
Related papers

Related papers: GaGA: Towards Interactive Global Geolocation Assis…

200 papers

Motivated by the goal of achieving long-term drift-free camera pose estimation in complex scenarios, we propose a global positioning framework fusing visual, inertial and Global Navigation Satellite System (GNSS) measurements in multiple…

Robotics · Computer Science 2022-01-06 Bing Han , Zhongyang Xiao , Shuai Huang , Tao Zhang

Cross-view geo-localization (CVGL) aims to estimate the geographic location of a street image by matching it with a corresponding aerial image. This is critical for autonomous navigation and mapping in complex real-world scenarios. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Songsong Ouyang , Yingying Zhu

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Azmine Toushik Wasi , Shahriyar Zaman Ridoy , Koushik Ahamed Tonmoy , Kinga Tshering , S. M. Muhtasimul Hasan , Wahid Faisal , Tasnim Mohiuddin , Md Rizwan Parvez

The standard approach for visual place recognition is to use global image descriptors to retrieve the most similar database images for a given query image. The results can then be further improved with re-ranking methods that re-order the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Gustav Hanning , Gabrielle Flood , Viktor Larsson

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Zixuan Song , Jing Zhang , Di Wang , Zidie Zhou , Wenbin Liu , Haonan Guo , En Wang , Bo Du

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Parth Parag Kulkarni , Rohit Gupta , Prakash Chandra Chhipa , Mubarak Shah

The task of motion forecasting is critical for self-driving vehicles (SDVs) to be able to plan a safe maneuver. Towards this goal, modern approaches reason about the map, the agents' past trajectories and their interactions in order to…

Robotics · Computer Science 2022-11-10 Alexander Cui , Sergio Casas , Kelvin Wong , Simon Suo , Raquel Urtasun

Recently, linear complexity sequence modeling networks have achieved modeling capabilities similar to Vision Transformers on a variety of computer vision tasks, while using fewer FLOPs and less memory. However, their advantage in terms of…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Bencheng Liao , Xinggang Wang , Lianghui Zhu , Qian Zhang , Chang Huang

Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, existing methods typically rely on complex architectures to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Zhengxian Wu , Chuanrui Zhang , Shenao Jiang , Hangrui Xu , Zirui Liao , Luyuan Zhang , Huaqiu Li , Peng Jiao , Haoqian Wang

We propose an image representation and matching approach that substantially improves visual-based location estimation for images. The main novelty of the approach, called distinctive visual element matching (DVEM), is its use of…

Multimedia · Computer Science 2016-01-29 Xinchao Li , Martha A. Larson , Alan Hanjalic

Geospatial Location Embedding (GLE) helps a Large Language Model (LLM) assimilate and analyze spatial data. GLE emergence in Geospatial Artificial Intelligence (GeoAI) is precipitated by the need for deeper geospatial awareness in our…

Information Retrieval · Computer Science 2024-01-22 Sean Tucker

Visual localization has traditionally been formulated as a pair-wise pose regression problem. Existing approaches mainly estimate relative poses between two images and employ a late-fusion strategy to obtain absolute pose estimates.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Tianchen Deng , Wenhua Wu , Kunzhen Wu , Guangming Wang , Siting Zhu , Shenghai Yuan , Xun Chen , Guole Shen , Zhe Liu , Hesheng Wang

The rapid advancement of multimodal large language models (LLMs) has opened new frontiers in artificial intelligence, enabling the integration of diverse large-scale data types such as text, images, and spatial information. In this paper,…

Artificial Intelligence · Computer Science 2025-03-21 Long Yuan , Fengran Mo , Kaiyu Huang , Wenjie Wang , Wangyuxuan Zhai , Xiaoyu Zhu , You Li , Jinan Xu , Jian-Yun Nie

Geometry mathematics problems pose significant challenges for large language models (LLMs) because they involve visual elements and spatial reasoning. Current methods primarily rely on symbolic character awareness to address these problems.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Shihao Xu , Yiyang Luo , Wei Shi

Simultaneous Localization and Mapping (SLAM) is one of the most important environment-perception and navigation algorithms for computer vision, robotics, and autonomous cars/drones. Hence, high quality and fast mapping becomes a fundamental…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Runfa Blark Li , Mahdi Shaghaghi , Keito Suzuki , Xinshuang Liu , Varun Moparthi , Bang Du , Walker Curtis , Martin Renschler , Ki Myung Brian Lee , Nikolay Atanasov , Truong Nguyen

Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Yikun Wang , Zuyan Liu , Ziyi Wang , Han Hu , Pengfei Liu , Yongming Rao

Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich…

Robotics · Computer Science 2025-08-14 Lin Sun , Bin Xie , Yingfei Liu , Hao Shi , Tiancai Wang , Jiale Cao

The ability to accomplish tasks via natural language instructions is one of the most efficient forms of interaction between humans and technology. This efficiency has been translated into practical applications with generative AI tools now…

Computational Engineering, Finance, and Science · Computer Science 2025-01-27 Yared W. Bekele

Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Linger Deng , Yuliang Liu , Wenwen Yu , Zujia Zhang , Jianzhong Ju , Zhenbo Luo , Xiang Bai

Answering real-world geospatial questions--such as finding restaurants along a travel route or amenities near a landmark--requires reasoning over both geographic relationships and semantic user intent. However, existing large language…

Information Retrieval · Computer Science 2025-06-12 Dazhou Yu , Riyang Bao , Ruiyu Ning , Jinghong Peng , Gengchen Mai , Liang Zhao
‹ Prev 1 4 5 6 7 8 10 Next ›