English
Related papers

Related papers: GEO-Bench-2: From Performance to Capability, Rethi…

200 papers

Generative AI based on foundation models provides a first glimpse into the world represented by machines trained on vast amounts of multimodal data ingested by these models during training. If we consider the resulting models as knowledge…

Computers and Society · Computer Science 2024-04-12 Zilong Liu , Krzysztof Janowicz , Kitty Currier , Meilin Shi

Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standardized evaluation metrics, datasets, and protocols. Existing systems are assessed using different…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Jiaming Wang , Diwen Liu , Jizhuo Chen , Harold Soh

In this paper, we introduce a benchmarking framework within the open-source NVIDIA PhysicsNeMo-CFD framework designed to systematically assess the accuracy, performance, scalability, and generalization capabilities of AI models for…

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yushuo Zheng , Jiangyong Ying , Huiyu Duan , Chunyi Li , Zicheng Zhang , Jing Liu , Xiaohong Liu , Guangtao Zhai

Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jiaxin Ge , Grace Luo , Heekyung Lee , Nishant Malpani , Long Lian , XuDong Wang , Aleksander Holynski , Trevor Darrell , Sewon Min , David M. Chan

Earth Observation (EO) is essential for perceiving dynamic land surface changes, yet deploying autonomous EO in open environments is hindered by the immense diversity of multi-source data and heterogeneous tasks. While remote sensing agents…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Sijie Zhao , Feng Liu , Xueliang Zhang , Hao Chen , Xinyu Gu , Zhe Jiang , Fenghua Ling , Ben Fei , Wenlong Zhang , Junjue Wang , Weihao Xuan , Pengfeng Xiao , Naoto Yokoya , Lei Bai

Comprehensive evaluation of geospatial foundation models (Geo-FMs) requires benchmarking across diverse tasks, sensors, and geographic regions. However, most existing benchmark datasets are limited to segmentation or classification tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Aaron Banze , Timothée Stassin , Nassim Ait Ali Braham , Rıdvan Salih Kuzu , Simon Besnard , Michael Schmitt

We present AiTLAS: Benchmark Arena -- an open-source benchmark suite for evaluating state-of-the-art deep learning approaches for image classification in Earth Observation (EO). To this end, we present a comprehensive comparative analysis…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Ivica Dimitrovski , Ivan Kitanovski , Dragi Kocev , Nikola Simidjievski

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models' decisions has grown…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Ruoyu Chen , Siyuan Liang , Jingzhi Li , Shiming Liu , Maosen Li , Zhen Huang , Hua Zhang , Xiaochun Cao

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Bohao Li , Yuying Ge , Yixiao Ge , Guangzhi Wang , Rui Wang , Ruimao Zhang , Ying Shan

Earth Observation (EO) data encompass a vast range of remotely sensed information, featuring multi-sensor and multi-temporal, playing an indispensable role in understanding our planet's dynamics. Recently, Vision Language Models (VLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Xizhe Xue , Xiao Xiang Zhu

Foundation models have the potential to transform the landscape of remote sensing (RS) data analysis by enabling large computer vision models to be pre-trained on vast amounts of remote sensing data. These models can then be fine-tuned with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Caleb S. Spradlin , Jordan A. Caraballo-Vega , Jian Li , Mark L. Carroll , Jie Gong , Paul M. Montesano

Tabular data, owing to its ubiquitous presence in real-world domains, has garnered significant attention in machine learning research. While tree-based models have long dominated tabular machine learning tasks, the recently proposed deep…

Machine Learning · Computer Science 2025-05-23 Zi-Jian Cheng , Zi-Yi Jia , Zhi Zhou , Yu-Feng Li , Lan-Zhe Guo

Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding. However, most methods rely on training…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Yue Zhou , Zhihang Zhong , Xue Yang

Recent progress in self-supervision shows that pre-training large neural networks on vast amounts of unsupervised data can lead to impressive increases in generalisation for downstream tasks. Such models, recently coined as foundation…

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Dingming Li , Hongxing Li , Zixuan Wang , Yuchen Yan , Hang Zhang , Siqi Chen , Guiyang Hou , Shengpei Jiang , Wenqi Zhang , Yongliang Shen , Weiming Lu , Yueting Zhuang

The growing availability of Earth Observation (EO) data and recent advances in Computer Vision have driven rapid progress in machine learning for EO, producing domain-specific models at ever-increasing scales. Yet this progress risks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Tasos Papazafeiropoulos , Nikolaos Ioannis Bountos , Nikolas Papadopoulos , Ioannis Papoutsis

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Juntong Wang , Jiarui Wang , Huiyu Duan , Jiaxiang Kang , Guangtao Zhai , Xiongkuo Min

Large foundation models (FMs) are transforming Earth science by integrating heterogeneous multimodal data, such as multi-platform imagery, gridded reanalysis data, diverse geophysical and geochemical observations, and domain-specific text,…

Instrumentation and Methods for Astrophysics · Physics 2026-05-14 Xiangyu Zhao , Bo Liu , Yuehan Zhang , Zelin Song , Wanghan Xu , Feng Liu , Fengxiang Wang , Ben Fei , Fenghua Ling , Wangxu Wei , Wenlong Zhang , Xiao-Ming Wu

Foundation models are transformative in artificial intelligence, but building them from scratch, especially for mobility trajectories, is not yet clear or documented. This tutorial bridges this gap by demonstrating the steps and code of a…

Artificial Intelligence · Computer Science 2025-11-26 Gaspard Merten , Mahmoud Sakr , Gilles Dejaegere