English
Related papers

Related papers: GeoGrid-Bench: Can Foundation Models Understand Mu…

200 papers

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Linger Deng , Yuliang Liu , Wenwen Yu , Zujia Zhang , Jianzhong Ju , Zhenbo Luo , Xiang Bai

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jonathan Roberts , Timo Lüddecke , Rehan Sheikh , Kai Han , Samuel Albanie

Integrating gridded weather and earth observation data into impact evaluations holds great promise. It allows researchers to capture environmental context, external shocks, and even to measure outcomes (e.g., land cover change, agricultural…

Physics and Society · Physics 2025-10-08 Elinor Benami , Mike Cecil , Anna Josephson , Gina Maskell , Jeffrey D. Michler

Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Pierre Adorni , Minh-Tan Pham , Stéphane May , Sébastien Lefèvre

As climate change intensifies, the urgency for accurate global-scale disaster predictions grows. This research presents a novel multimodal disaster prediction framework, combining weather statistics, satellite imagery, and textual insights.…

Machine Learning · Computer Science 2023-10-02 Gengyin Liu , Huaiyang Zhong

Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, and other high-stakes Earth-observation tasks. Yet the published work about these models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Isaac Corley , Nils Lehmann , Caleb Robinson , Gabriel Tseng , Anthony Fuller , Hamed Alemohammad , Evan Shelhamer , Jennifer Marcus , Hannah Kerner

Current foundation models exhibit impressive capabilities when prompted either with text only or with both image and text inputs. But do their capabilities change depending on the input modality? In this work, we propose…

Artificial Intelligence · Computer Science 2024-08-20 Deqing Fu , Ruohao Guo , Ghazal Khalighinejad , Ollie Liu , Bhuwan Dhingra , Dani Yogatama , Robin Jia , Willie Neiswanger

Foundation models, i.e., very large deep learning models, have demonstrated impressive performances in various language and vision tasks that are otherwise difficult to reach using smaller-size models. The major success of GPT-type of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Yiqun Xie , Zhihao Wang , Weiye Chen , Zhili Li , Xiaowei Jia , Yanhua Li , Ruichen Wang , Kangyang Chai , Ruohan Li , Sergii Skakun

Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have been limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Baichuan Zhou , Haote Yang , Dairong Chen , Junyan Ye , Tianyi Bai , Jinhua Yu , Songyang Zhang , Dahua Lin , Conghui He , Weijia Li

This report focuses on spatial data intelligent large models, delving into the principles, methods, and cutting-edge applications of these models. It provides an in-depth discussion on the definition, development history, current status,…

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

Change detection, as an important and widely applied technique in the field of remote sensing, aims to analyze changes in surface areas over time and has broad applications in areas such as environmental monitoring, urban development, and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Zihan Yu , Tianxiao Li , Yuxin Zhu , Rongze Pan

AI-assisted coding has rapidly reshaped software practice and research workflows, yet today's models still struggle to produce correct code for complex 3D geometric vision. If models could reliably write such code, the research of our…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Wenyi Li , Renkai Luo , Yue Yu , Huan-ang Gao , Mingju Gao , Li Yuan , Chaoyou Fu , Hao Zhao

With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focus on objective…

Computation and Language · Computer Science 2025-09-24 Haokun Li , Yazhou Zhang , Jizhi Ding , Qiuchi Li , Peng Zhang

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning…

Computation and Language · Computer Science 2025-04-25 Liangyu Xu , Yingxiu Zhao , Jingyun Wang , Yingyao Wang , Bu Pi , Chen Wang , Mingliang Zhang , Jihao Gu , Xiang Li , Xiaoyong Zhu , Jun Song , Bo Zheng

Central to Earth observation is the trade-off between spatial and temporal resolution. For temperature, this is especially critical because real-world applications require high spatiotemporal resolution data. Current technology allows for…

Image and Video Processing · Electrical Eng. & Systems 2025-07-15 Shengjie Liu , Lu Zhang , Siqin Wang

Evaluating foundation models under appropriate adaptation settings is essential for understanding the quality and transferability of the learned representations. Recent EEG foundation models have demonstrated promising transfer capabilities…

Machine Learning · Computer Science 2026-05-28 Aditya Kommineni , Emily Zhou , Kleanthis Avramidis , Tiantian Feng , Shrikanth Narayanan

We introduce SciVer, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context. SciVer consists of 3,000 expert-annotated examples over 1,113 scientific…

Computation and Language · Computer Science 2025-06-19 Chengye Wang , Yifei Shen , Zexi Kuang , Arman Cohan , Yilun Zhao

Recent advances in large language models (LLMs) have fueled growing interest in automating geospatial analysis and GIS workflows, yet their actual capabilities remain uncertain. In this work, we call for rigorous evaluation of LLMs on…

Software Engineering · Computer Science 2025-09-09 Qianheng Zhang , Song Gao , Chen Wei , Yibo Zhao , Ying Nie , Ziru Chen , Shijie Chen , Yu Su , Huan Sun