English
Related papers

Related papers: Mind over Space: Can Multimodal Large Language Mod…

200 papers

Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Sicheng Feng , Song Wang , Shuyi Ouyang , Lingdong Kong , Zikai Song , Jianke Zhu , Huan Wang , Xinchao Wang

Extracting concepts and understanding relationships from videos is essential in Video-Based Design (VBD), where videos serve as a primary medium for exploration but require significant effort in managing meta-information. Mind maps, with…

Human-Computer Interaction · Computer Science 2025-01-17 Tianhao He , Karthi Saravanan , Evangelos Niforatos , Gerd Kortuem

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

Top-view perspective denotes a typical way in which humans read and reason over different types of maps, and it is vital for localization and navigation of humans as well as of `non-human' agents, such as the ones backed by large…

Computation and Language · Computer Science 2024-06-05 Chengzu Li , Caiqi Zhang , Han Zhou , Nigel Collier , Anna Korhonen , Ivan Vulić

Vision language models (VLMs) are an exciting emerging class of language models (LMs) that have merged classic LM capabilities with those of image processing systems. However, the ways that these capabilities combine are not always…

Computation and Language · Computer Science 2024-07-03 Qiucheng Wu , Handong Zhao , Michael Saxon , Trung Bui , William Yang Wang , Yang Zhang , Shiyu Chang

Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these…

Artificial Intelligence · Computer Science 2026-03-03 Qianqian Bai , Zhongpu Chen , Ling Luo , Huaming Du , Yuqian Lei , Ziyun Jiao

Mental visualization, the ability to construct and manipulate visual representations internally, is a core component of human cognition and plays a vital role in tasks involving reasoning, prediction, and abstraction. Despite the rapid…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Mohammad Shahab Sepehri , Berk Tinaz , Zalan Fabian , Mahdi Soltanolkotabi

Object goal navigation (ObjectNav) is a fundamental task in embodied AI, requiring an agent to locate a target object in previously unseen environments. This task is particularly challenging because it requires both perceptual and cognitive…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Yihan Cao , Jiazhao Zhang , Zhinan Yu , Shuzhen Liu , Zheng Qin , Qin Zou , Bo Du , Kai Xu

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Artificial Intelligence · Computer Science 2026-05-12 Himanshu Gupta , Shreyas Verma , Ujjwala Anantheswaran , Kevin Scaria , Mihir Parmar , Swaroop Mishra , Chitta Baral

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a benchmark to evaluate whether video-large language models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Baining Zhao , Jianjie Fang , Zichao Dai , Ziyou Wang , Jirong Zha , Weichen Zhang , Chen Gao , Yue Wang , Jinqiang Cui , Xinlei Chen , Yong Li

Benchmarking spatial reasoning in multimodal large language models (MLLMs) has attracted growing interest in computer vision due to its importance for embodied AI and other agentic systems that require precise interaction with the physical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Zelin Xu , Yupu Zhang , Saugat Adhikari , Saiful Islam , Tingsong Xiao , Zibo Liu , Shigang Chen , Da Yan , Zhe Jiang

Unlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on explicit 3D inputs or specialized model architectures, we ask:…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Fangrui Zhu , Hanhui Wang , Yiming Xie , Jing Gu , Tianye Ding , Jianwei Yang , Huaizu Jiang

Spatial visualization is the mental ability to imagine, transform, and manipulate the spatial characteristics of objects and actions. This intelligence is a part of human cognition where actions and perception are connected on a mental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Nilay Yilmaz , Maitreya Patel , Naga Sai Abhiram Kusumba , Yixuan He , Yezhou Yang

Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently operate planning within the language space, lacking the vision…

Computer Vision and Pattern Recognition · Computer Science 2024-02-19 Jun Cen , Chenfei Wu , Xiao Liu , Shengming Yin , Yixuan Pei , Jinglong Yang , Qifeng Chen , Nan Duan , Jianguo Zhang

Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Siting Wang , Minnan Pei , Luoyang Sun , Cheng Deng , Yuchen Li , Kun Shao , Zheng Tian , Haifeng Zhang , Jun Wang

The advent of Multimodal Large Language Models, leveraging the power of Large Language Models, has recently demonstrated superior multimodal understanding and reasoning abilities, heralding a new era for artificial general intelligence.…

Artificial Intelligence · Computer Science 2025-04-14 Lu Qiu , Yi Chen , Yuying Ge , Yixiao Ge , Ying Shan , Xihui Liu

The Visual-and-Language Navigation (VLN) task requires understanding a textual instruction to navigate a natural indoor environment using only visual information. While this is a trivial task for most humans, it is still an open problem for…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Joaquin Ossandón , Benjamin Earle , Álvaro Soto

Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for…

Cognitive maps are a proposed concept on how the brain efficiently organizes memories and retrieves context out of them. The entorhinal-hippocampal complex is heavily involved in episodic and relational memory processing, as well as spatial…

Neurons and Cognition · Quantitative Biology 2024-01-04 Paul Stoewer , Achim Schilling , Andreas Maier , Patrick Krauss

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Zeyuan Yang , Xueyang Yu , Delin Chen , Maohao Shen , Chuang Gan