English
Related papers

Related papers: SpatialBench: Is Your Spatial Foundation Model an …

200 papers

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Sashuai Zhou , Qiang Zhou , Junpeng Ma , Yue Cao , Ruofan Hu , Ziang Zhang , Xiaoda Yang , Zhibin Wang , Jun Song , Cheng Yu , Bo Zheng , Zhou Zhao

Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a comprehensive framework spanning classification, segmentation, regression, object…

Existing benchmarks for multimodal learning in Earth science offer limited, siloed coverage of Earth's spheres and their cross-sphere interactions, typically restricting evaluation to the human-activity sphere of atmosphere and to at most…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Fengxiang Wang , Mingshuo Chen , Xuming He , Yi-Fan Zhang , Yueying Li , Feng Liu , Zijie Guo , Zhenghao Hu , Jiong Wang , Jingyi Xu , Zhangrui Li , Junchao Gong , Di Wang , Fenghua Ling , Ben Fei , Weijia Li , Long Lan , Wenjing Yang

Stochastic Partial Differential Equations (SPDEs) driven by random noise play a central role in modeling physical processes with rough spatio-temporal dynamics, such as turbulence flows, superconductors, and quantum dynamics. Although…

Machine Learning · Computer Science 2026-05-18 Yuantu Zhu , Zheyan Li , Dai Shi , Luke Thompson , Oliver Nash , Jose Miguel Lara Rangel , Siran Li , Bingguang Chen , Rongchan Zhu , Qi Meng , Hao Ni

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Bobby Azad , Reza Azad , Sania Eskandari , Afshin Bozorgpour , Amirhossein Kazerouni , Islem Rekik , Dorit Merhof

Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial information expressed in procedural code remains…

Artificial Intelligence · Computer Science 2026-02-11 Shixian Luo , Zezhou Zhu , Yu Yuan , Yuncheng Yang , Lianlei Shan , Yong Wu

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear…

Artificial Intelligence · Computer Science 2025-04-21 Santhosh Kumar Ramakrishnan , Erik Wijmans , Philipp Kraehenbuehl , Vladlen Koltun

Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot setting, requiring full solution generation in a single response,…

Artificial Intelligence · Computer Science 2026-04-13 Lars Benedikt Kaesberg , Tianyu Yang , Niklas Bauer , Terry Ruas , Jan Philip Wahle , Bela Gipp

Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making under uncertainty. However, common existing benchmarks…

Artificial Intelligence · Computer Science 2026-05-12 Wenjie Tang , Yuan Zhou , Erqiang Xu , Keyan Cheng , Minne Li , Liquan Xiao

It is unclear whether strong forecasting performance reflects genuine temporal understanding or the ability to reason under contextual and event-driven conditions. We introduce TemporalBench, a multi-domain benchmark designed to evaluate…

Artificial Intelligence · Computer Science 2026-02-17 Muyan Weng , Defu Cao , Wei Yang , Yashaswi Sharma , Yan Liu

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Rishi Upadhyay , Howard Zhang , Jim Solomon , Ayush Agrawal , Pranay Boreddy , Shruti Satya Narayana , Yunhao Ba , Alex Wong , Celso M de Melo , Achuta Kadambi

Accurate epidemic forecasting is crucial for public health response, resource allocation, and outbreak intervention, but remains difficult with sparse, noisy, and highly non-stationary data. Because epidemics unfold across interacting…

Artificial Intelligence · Computer Science 2026-05-08 Ruiqi Lyu , Alistair Turcan , Bryan Wilder

The rapid evolution of large language models (LLMs) holds promise for reforming the methodology of spatio-temporal data mining. However, current works for evaluating the spatio-temporal understanding capability of LLMs are somewhat limited…

Computation and Language · Computer Science 2024-06-28 Wenbin Li , Di Yao , Ruibo Zhao , Wenjie Chen , Zijie Xu , Chengxue Luo , Chang Gong , Quanliang Jing , Haining Tan , Jingping Bi

Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Pukun Zhao , Longxiang Wang , Miaowei Wang , Chen Chen , Fanqing Zhou , Haojian Huang

Data-driven modeling of fluid dynamics has advanced rapidly with neural PDE solvers, yet a fair and strong benchmark remains fragmented due to the absence of unified PDE datasets and standardized evaluation protocols. Although architectural…

Fluid Dynamics · Physics 2026-05-22 Haixin Wang , Ruoyan Li , Fred Xu , Fang Sun , Kaiqiao Han , Zijie Huang , Ching Chang , Xiao Luo , Wei Wang , Yizhou Sun

Autoscaling has become a baseline expectation for cloud-native big data processing, and the design space has expanded beyond rule-based heuristics to include learned controllers and, most recently, large language model (LLM) agents. Yet…

Information Retrieval · Computer Science 2026-05-13 Venkata Krishna Prasanth Budigi , Siri Chandana Sirigiri

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiyuan Yan , Yong Zhang , Xinhang Yuan , Siwei Lyu , Baoyuan Wu

Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is…

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A comprehensive evaluation benchmark for long-form speech is…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-28 Changhao Pan , Rui Yang , Han Wang , Zhuan Zhou , Xuming He , Wenxiang Guo , Ziyue Jiang , Ruiqi Li , Yu Zhang , Chenyuhao Wen , Ke Lei , Xiang Yin , Jingyu Lu , Zhiyuan Zhu , Zhou Zhao