English
Related papers

Related papers: Everything in Its Place: Benchmarking Spatial Inte…

200 papers

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Juntong Wang , Jiarui Wang , Huiyu Duan , Jiaxiang Kang , Guangtao Zhai , Xiongkuo Min

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal…

As Large Language Models (LLMs) increasingly power autonomous agents in robotics and embodied AI, understanding their spatial reasoning capabilities becomes crucial for ensuring reliable real-world deployment. Despite advances in language…

Artificial Intelligence · Computer Science 2025-07-29 Hafsteinn Einarsson

Spatial intelligence (SI) represents a cognitive ability encompassing the visualization, manipulation, and reasoning about spatial relationships, underpinning disciplines from neuroscience to robotics. We introduce SITE, a benchmark dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Wenqi Wang , Reuben Tan , Pengyue Zhu , Jianwei Yang , Zhengyuan Yang , Lijuan Wang , Andrey Kolobov , Jianfeng Gao , Boqing Gong

Automating Text-to-Image (T2I) model evaluation is challenging; a judge model must be used to score correctness, and test prompts must be selected to be challenging for current T2I models but not the judge. We argue that satisfying these…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Amita Kamath , Kai-Wei Chang , Ranjay Krishna , Luke Zettlemoyer , Yushi Hu , Marjan Ghazvininejad

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jiarui Wang , Huiyu Duan , Yu Zhao , Juntong Wang , Guangtao Zhai , Xiongkuo Min

We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through the task of converting apartment photographs into accurate 2D floor plans. While the input modality (photographs) is well…

Artificial Intelligence · Computer Science 2025-10-01 Lukas Petersson , Axel Backlund , Axel Wennstöm , Hanna Petersson , Callum Sharrock , Arash Dabiri

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks…

Artificial Intelligence · Computer Science 2026-05-08 Peiran Xu , Sudong Wang , Yao Zhu , Jianing Li , Gege Qi , Yunjian Zhang

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Wufei Ma , Haoyu Chen , Guofeng Zhang , Yu-Cheng Chou , Jieneng Chen , Celso M de Melo , Alan Yuille

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical…

Machine Learning · Computer Science 2025-02-07 Ivana Kajić , Olivia Wiles , Isabela Albuquerque , Matthias Bauer , Su Wang , Jordi Pont-Tuset , Aida Nematzadeh

Recent text-to-image (T2I) generation models have achieved remarkable sucess by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data quantity over availability (closed vs open source) and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 L. Degeorge , A. Ghosh , N. Dufour , D. Picard , V. Kalogeiton

Spatial understanding over continuous visual input is crucial for MLLMs to evolve into general-purpose assistants in physical environments. Yet there is still no comprehensive benchmark that holistically assesses the progress toward this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Jingli Lin , Runsen Xu , Shaohao Zhu , Sihan Yang , Peizhou Cao , Yunlong Ran , Miao Hu , Chenming Zhu , Yiman Xie , Yilin Long , Wenbo Hu , Dahua Lin , Tai Wang , Jiangmiao Pang

Environment designers in the entertainment industry create imaginative 2D and 3D scenes for games, films, and television, requiring both fine-grained control of specific details and consistent global coherence. Designers have increasingly…

Human-Computer Interaction · Computer Science 2025-09-03 Wen-Fan Wang , Ting-Ying Lee , Chien-Ting Lu , Che-Wei Hsu , Nil Ponsa Campanyà , Yu Chen , Mike Y. Chen , Bing-Yu Chen

While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language alignment still struggle with the fine-grained semantic matching. Current approaches based…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Zijian Zhang , Xuhui Zheng , Xuecheng Wu , Chong Peng , Xuezhi Cao

The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Muna Numan Said , Aarib Zaidi , Rabia Usman , Sonia Okon , Praneeth Medepalli , Kevin Zhu , Vasu Sharma , Sean O'Brien

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jiayang Li , Shuo Cao , Xiaohui Li , Zhizhen Zhang , Kaiwen Zhu , Yule Duan , Yu Qiao , Jian Zhang , Yihao Liu

The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts -- where missed cues can stereotype communities and undermine…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Shravan Nayak , Mehar Bhatia , Xiaofeng Zhang , Verena Rieser , Lisa Anne Hendricks , Sjoerd van Steenkiste , Yash Goyal , Karolina Stańczak , Aishwarya Agrawal

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Song Tang , Peihao Gong , Kunyu Li , Kai Guo , Boyu Wang , Mao Ye , Jianwei Zhang , Xiatian Zhu

Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically apply uniform annotation mechanisms, such as Likert-scale or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Abdelrahman Eldesokey , Merey Ramazanova , Ahmad Sait , Ansar Khangeldin , Karen Sanchez , Tong Zhang , Bernard Ghanem