English
Related papers

Related papers: Exploring Spatial Intelligence from a Generative P…

200 papers

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Hongxing Li , Dingming Li , Zixuan Wang , Yuchen Yan , Hang Wu , Wenqi Zhang , Yongliang Shen , Weiming Lu , Jun Xiao , Yueting Zhuang

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Mingxin Liu , Ziqian Fan , Zhaokai Wang , Leyao Gu , Zirun Zhu , Yiguo He , Yuchen Yang , Changyao Tian , Xiangyu Zhao , Ning Liao , Shaofeng Zhang , Qibing Ren , Zhihang Zhong , Xuanhe Zhou , Junchi Yan , Xue Yang

In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent's representations during training or via use as part of an explicit planning…

The rapidly developing field of large multimodal models (LMMs) has led to the emergence of diverse models with remarkable capabilities. However, existing benchmarks fail to comprehensively, objectively and accurately evaluate whether LMMs…

The rapid advancement of generative artificial intelligence has enabled models capable of producing complex textual and visual outputs; however, their decision-making processes remain largely opaque, limiting trust and accountability in…

Artificial Intelligence · Computer Science 2026-02-03 Zeinab Dehghani

Spatial competence is the quality of maintaining a consistent internal representation of an environment and using it to infer discrete structure and plan actions under constraints. Prevailing spatial evaluations for large models are limited…

Artificial Intelligence · Computer Science 2026-04-14 Jash Vira , Ashley Harris

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear…

Artificial Intelligence · Computer Science 2025-04-21 Santhosh Kumar Ramakrishnan , Erik Wijmans , Philipp Kraehenbuehl , Vladlen Koltun

Generative AI offers new opportunities for automating urban planning by creating site-specific urban layouts and enabling flexible design exploration. However, existing approaches often struggle to produce realistic and practical designs at…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Qingyi Wang , Yuebing Liang , Yunhan Zheng , Kaiyuan Xu , Jinhua Zhao , Shenhao Wang

Over the past year, the development of large language models (LLMs) has brought spatial intelligence into focus, with much attention on vision-based embodied intelligence. However, spatial intelligence spans a broader range of disciplines…

We propose a new probabilistic framework that allows mobile robots to autonomously learn deep, generative models of their environments that span multiple levels of abstraction. Unlike traditional approaches that combine engineered models…

Robotics · Computer Science 2018-01-01 Andrzej Pronobis , Rajesh P. N. Rao

Multimodal large language models (MLLMs) have made significant progress in integrating visual and linguistic understanding. Existing benchmarks typically focus on high-level semantic capabilities, such as scene understanding and visual…

Computation and Language · Computer Science 2025-02-18 Shangyu Xing , Changhao Xiang , Yuteng Han , Yifan Yue , Zhen Wu , Xinyu Liu , Zhangtai Wu , Fei Zhao , Xinyu Dai

Self-supervised learning is a popular and powerful method for utilizing large amounts of unlabeled data, for which a wide variety of training objectives have been proposed in the literature. In this study, we perform a Bayesian analysis of…

Machine Learning · Computer Science 2023-02-08 Emanuele Sansone , Robin Manhaeve

High spectral dimensionality and the shortage of annotations make hyperspectral image (HSI) classification a challenging problem. Recent studies suggest that convolutional neural networks can learn discriminative spatial features, which…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Zilong Zhong , Jonathan Li

Recent advances in agentic AI have led to systems capable of autonomous task execution and language-based reasoning, yet their spatial reasoning abilities remain limited and underexplored, largely constrained to symbolic and sequential…

Artificial Intelligence · Computer Science 2025-09-12 Bui Duc Manh , Soumyaratna Debnath , Zetong Zhang , Shriram Damodaran , Arvind Kumar , Yueyi Zhang , Lu Mi , Erik Cambria , Lin Wang

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in spatially-aware tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Haiyi Qiu , Kaihang Pan , Jiacheng Li , Juncheng Li , Siliang Tang , Yueting Zhuang

Satellite imagery and remote sensing provide explanatory variables at relatively high resolutions for modeling geospatial phenomena, yet regional summaries are often desirable for analysis and actionable insight. In this paper, we propose a…

Machine Learning · Statistics 2017-12-15 Sam Kriegman , Marcin Szubert , Josh C. Bongard , Christian Skalka

Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For robot learning to…

Robotics · Computer Science 2024-09-04 Zoey Chen , Zhao Mandi , Homanga Bharadhwaj , Mohit Sharma , Shuran Song , Abhishek Gupta , Vikash Kumar

Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particularly in human image generation. Sociological research has…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hanjun Luo , Haoyu Huang , Ziye Deng , Xinfeng Li , Hewei Wang , Yingbin Jin , Yang Liu , Wenyuan Xu , Zuozhu Liu

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Sizhong Qin , Ramon Elias Weber , Xinzheng Lu
‹ Prev 1 3 4 5 6 7 10 Next ›