English
Related papers

Related papers: MapVerse: A Benchmark for Geospatial Question Answ…

200 papers

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Muhammad Sohail Danish , Muhammad Akhtar Munir , Syed Roshaan Ali Shah , Kartik Kuckreja , Fahad Shahbaz Khan , Paolo Fraccaro , Alexandre Lacoste , Salman Khan

Large Vision Language Models (LVLMs) have demonstrated remarkable abilities in understanding and reasoning about both visual and textual information. However, existing evaluation methods for LVLMs, primarily based on benchmarks like Visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Xinyu Wang , Bohan Zhuang , Qi Wu

As Large Language Models (LLMs) increasingly power autonomous agents in robotics and embodied AI, understanding their spatial reasoning capabilities becomes crucial for ensuring reliable real-world deployment. Despite advances in language…

Artificial Intelligence · Computer Science 2025-07-29 Hafsteinn Einarsson

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuan Feng , Yue Yang , Xiaohan He , Jiatong Zhao , Jianlong Chen , Zijun Chen , Daocheng Fu , Qi Liu , Renqiu Xia , Bo Zhang , Junchi Yan

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Azmine Toushik Wasi , Shahriyar Zaman Ridoy , Koushik Ahamed Tonmoy , Kinga Tshering , S. M. Muhtasimul Hasan , Wahid Faisal , Tasnim Mohiuddin , Md Rizwan Parvez

Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior research in this domain has been constrained by non-quantifiable…

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Jihan Yang , Shusheng Yang , Anjali W. Gupta , Rilyn Han , Li Fei-Fei , Saining Xie

Puzzles have long served as compact and revealing probes of human cognition, isolating abstraction, rule discovery, and systematic reasoning with minimal reliance on prior knowledge. Leveraging these properties, visual puzzles have recently…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Maria Lymperaiou , Vasileios Karampinis , Giorgos Filandrianos , Angelos Vlachos , Chrysoula Zerva , Athanasios Voulodimos

The advancement of large language models (LLMs) has significantly broadened the scope of applications in natural language processing, with multi-modal LLMs extending these capabilities to integrate and interpret visual data. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Bingchen Zhao , Yongshuo Zong , Letian Zhang , Timothy Hospedales

Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Pan Lu , Hritik Bansal , Tony Xia , Jiacheng Liu , Chunyuan Li , Hannaneh Hajishirzi , Hao Cheng , Kai-Wei Chang , Michel Galley , Jianfeng Gao

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs. To study the…

Computation and Language · Computer Science 2025-06-09 Yingjie Zhu , Xuefeng Bai , Kehai Chen , Yang Xiang , Jun Yu , Min Zhang

Earth vision has achieved milestones in geospatial object recognition but lacks exploration in object-relational reasoning, limiting comprehensive scene understanding. To address this, a progressive Earth vision-language understanding and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Junjue Wang , Yanfei Zhong , Zihang Chen , Zhuo Zheng , Ailong Ma , Liangpei Zhang

Spatial reasoning is a core component of human cognition, enabling individuals to perceive, comprehend, and interact with the physical world. It relies on a nuanced understanding of spatial structures and inter-object relationships, serving…

Artificial Intelligence · Computer Science 2025-08-27 Zesen Lyu , Dandan Zhang , Wei Ye , Fangdi Li , Zhihang Jiang , Yao Yang

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yolo Y. Tang , Pinxin Liu , Zhangyun Tan , Mingqian Feng , Rui Mao , Chao Huang , Jing Bi , Yunzhong Xiao , Susan Liang , Hang Hua , Ali Vosoughi , Luchuan Song , Zeliang Zhang , Chenliang Xu

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Haotian Xu , Yue Hu , Zhengqiu Zhu , Chen Gao , Ziyou Wang , Junreng Rao , Wenhao Lu , Weishi Li , Quanjun Yin , Yong Li

The ability to process information from multiple modalities and to reason through it step-by-step remains a critical challenge in advancing artificial intelligence. However, existing reasoning benchmarks focus on text-only reasoning, or…

Artificial Intelligence · Computer Science 2025-07-01 Yulun Jiang , Yekun Chai , Maria Brbić , Michael Moor

Large Language Models (LLM) have emerged as a tool for robots to generate task plans using common sense reasoning. For the LLM to generate actionable plans, scene context must be provided, often through a map. Recent works have shifted from…

Robotics · Computer Science 2024-09-25 Mike Zhang , Kaixian Qu , Vaishakh Patil , Cesar Cadena , Marco Hutter

We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs covering mostly relative spatial understanding and the order of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Nils Hoehing , Mayug Maniparambil , Ellen Rushe , Noel E. O'Connor , Anthony Ventresque

Spatial relations are a basic part of human cognition. However, they are expressed in natural language in a variety of ways, and previous work has suggested that current vision-and-language models (VLMs) struggle to capture relational…

Computation and Language · Computer Science 2023-03-23 Fangyu Liu , Guy Emerson , Nigel Collier