English
Related papers

Related papers: From Pixels to Places: A Systematic Benchmark for …

200 papers

This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs). Leveraging 1.46 million Mapillary street-level images, GeoChain pairs each image with a…

Artificial Intelligence · Computer Science 2025-09-10 Sahiti Yerramilli , Nilay Pande , Rynaa Grover , Jayant Sravan Tamarapalli

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Hang Du , Jiayang Zhang , Guoshun Nan , Wendi Deng , Zhenyan Chen , Chenyang Zhang , Wang Xiao , Shan Huang , Yuqi Pan , Tao Qi , Sicong Leng

The advances in Vision-Language models (VLMs) offer exciting opportunities for robotic applications involving image geo-localization, the problem of identifying the geo-coordinates of a place based on visual data only. Recent research works…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Sania Waheed , Bruno Ferrarini , Michael Milford , Sarvapali D. Ramchurn , Shoaib Ehsan

Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a two-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Pengyue Jia , Seongheon Park , Song Gao , Xiangyu Zhao , Sharon Li

Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visual question answering, and visual grounding. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Ruizhe Ou , Yuan Hu , Fan Zhang , Jiaxin Chen , Yu Liu

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yushuo Zheng , Jiangyong Ying , Huiyu Duan , Chunyi Li , Zicheng Zhang , Jing Liu , Xiaohong Liu , Guangtao Zhai

Large language models (LLMs) exhibit a variety of promising capabilities in robotics, including long-horizon planning and commonsense reasoning. However, their performance in place recognition is still underexplored. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Zonglin Lyu , Juexiao Zhang , Mingxuan Lu , Yiming Li , Chen Feng

Geospatial predictions are crucial for diverse fields such as disaster management, urban planning, and public health. Traditional machine learning methods often face limitations when handling unstructured or multi-modal data like street…

Computation and Language · Computer Science 2024-11-25 Zongrong Li , Junhao Xu , Siqin Wang , Yifan Wu , Haiyang Li

Multimodal large language models (MLLMs) have made significant progress in integrating visual and linguistic understanding. Existing benchmarks typically focus on high-level semantic capabilities, such as scene understanding and visual…

Computation and Language · Computer Science 2025-02-18 Shangyu Xing , Changhao Xiang , Yuteng Han , Yifan Yue , Zhen Wu , Xinyu Liu , Zhangtai Wu , Fei Zhao , Xinyu Dai

The ability to locate an object in an image according to natural language instructions is crucial for many real-world applications. In this work we propose LocateBench, a high-quality benchmark dedicated to evaluating this ability. We…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Ting-Rui Chiang , Joshua Robinson , Xinyan Velocity Yu , Dani Yogatama

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Existing spatial-temporal benchmarks mainly focus on egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Qinghongbing Xie , Zhaoyuan Xia , Feng Zhu , Lijun Gong , Ziyue Li , Rui Zhao , Long Zeng

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks that describe the salient and class-discriminative visual…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Grace Luo , Giscard Biamby , Trevor Darrell , Daniel Fried , Anna Rohrbach

The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yun Li , Yiming Zhang , Tao Lin , Xiangrui Liu , Wenxiao Cai , Zheng Liu , Bo Zhao

Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such as cropping, editing, or enhancement to uncover salient…

Street-level geolocalization from images is crucial for a wide range of essential applications and services, such as navigation, location-based recommendations, and urban planning. With the growing popularity of social media data and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yunus Serhat Bicakci , Joseph Shingleton , Anahid Basiri

Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. However, existing benchmarks primarily evaluate the embodied…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Haotian Xue , Yunhao Ge , Yu Zeng , Zhaoshuo Li , Ming-Yu Liu , Yongxin Chen , Jiaojiao Fan

In human reading and communication, individuals tend to engage in geospatial reasoning, which involves recognizing geographic entities and making informed inferences about their interrelationships. To mimic such cognitive process, current…

Computation and Language · Computer Science 2024-08-22 Yibo Yan , Joey Lee

Determining the precise geographic location of an image at a global scale remains an unsolved challenge. Standard image retrieval techniques are inefficient due to the sheer volume of images (>100M) and fail when coverage is insufficient.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Philipp Lindenberger , Paul-Edouard Sarlin , Jan Hosang , Matteo Balice , Marc Pollefeys , Simon Lynen , Eduard Trulls

Large language models have achieved great success in recent years, so as their variants in vision. Existing vision-language models can describe images in natural languages, answer visual-related questions, or perform complex reasoning about…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Jiarui Xu , Xingyi Zhou , Shen Yan , Xiuye Gu , Anurag Arnab , Chen Sun , Xiaolong Wang , Cordelia Schmid