中文
相关论文

相关论文: Why Do MLLMs Struggle with Spatial Understanding? …

200 篇论文

While diffusion models have shown exceptional capabilities in aesthetic image synthesis, they often struggle with complex spatial understanding and reasoning. Existing approaches resort to Multimodal Large Language Models (MLLMs) to enhance…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Wei Chen , Yancheng Long , Mingqiao Liu , Haojie Ding , Yankai Yang , Hongyang Wei , Yi-Fan Zhang , Bin Wen , Fan Yang , Tingting Gao , Han Li , Long Chen

Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large Language Models (LLMs) improve their reasoning…

计算与语言 · 计算机科学 2026-05-20 Husnain Amjad , Raja Khurram Shahzad , Aamir Shahzad , Mehwish Fatima

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle to reason about 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Om Khangaonkar , Hadi J. Rad , Hamed Pirsiavash

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. However, their performance is still lacking in physical domains that require…

人工智能 · 计算机科学 2025-07-04 Erle Zhu , Yadi Liu , Zhe Zhang , Xujun Li , Jin Zhou , Xinjie Yu , Minlie Huang , Hongning Wang

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Tanner Muturi , Blessing Agyei Kyem , Joshua Kofi Asamoah , Neema Jakisa Owor , Richard Dyzinela , Andrews Danyo , Yaw Adu-Gyamfi , Armstrong Aboah

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incoherent integration when acquiring information (Perception) and…

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Kevin Qu , Haozhe Qi , Mihai Dusmanu , Mahdi Rad , Rui Wang , Marc Pollefeys

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). Current methods predominantly rely on verbal descriptive…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Meng Cao , Haokun Lin , Haoyuan Li , Haoran Tang , Rongtao Xu , Dong An , Xue Liu , Ian Reid , Xiaodan Liang

Large language models (LLMs) demonstrate extraordinary abilities in a wide range of natural language processing (NLP) tasks. In this paper, we show that, beyond text understanding capability, LLMs are capable of processing text layouts that…

计算与语言 · 计算机科学 2024-08-29 Weiming Li , Manni Duan , Dong An , Yan Shao

Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, often failing at fundamental geometric tasks such as depth…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Zishan Liu , Ruoxi Zang , Yanglin Zhang , Wei Liu , Yin Zhang , Jian Yao , Jiayin Zheng , Zhengzhe Liu

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language…

机器人学 · 计算机科学 2025-11-03 Simindokht Jahangard , Mehrzad Mohammadi , Abhinav Dhall , Hamid Rezatofighi

Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversities between the natural and remote sensing (RS) images, the…

计算机视觉与模式识别 · 计算机科学 2024-03-11 Wei Zhang , Miaoxin Cai , Tong Zhang , Yin Zhuang , Xuerui Mao

Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simple yet revealing experiment, we interestingly find that language-only models, when provided…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yuting Li , Lai Wei , Kaipeng Zheng , Jingyuan Huang , Guilin Li , Bo Wang , Linghe Kong , Lichao Sun , Weiran Huang

This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to…

声音 · 计算机科学 2024-06-17 Changli Tang , Wenyi Yu , Guangzhi Sun , Xianzhao Chen , Tian Tan , Wei Li , Jun Zhang , Lu Lu , Zejun Ma , Yuxuan Wang , Chao Zhang

The mainstream paradigm of remote sensing image interpretation has long been dominated by vision-centered models, which rely on visual features for semantic understanding. However, these models face inherent limitations in handling…

人工智能 · 计算机科学 2026-01-28 Haifeng Li , Wang Guo , Haiyang Wu , Mengwei Wu , Jipeng Zhang , Qing Zhu , Yu Liu , Xin Huang , Chao Tao

Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Ruiheng Liu , Haihong Hao , Mingfei Han , Xin Gu , Kecheng Zhang , Changlin Li , Xiaojun Chang

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) in 360-degree…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Yuangong Chen , Wai Keung Wong , Jiaxing Li , Ioannis Patras , Xu Zheng

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Gui Wang , YongSong Zhou , Kaijun Deng , Wooi Ping Cheah , Rong Qu , Jianfeng Ren , Linlin Shen

Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research--relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive…

计算与语言 · 计算机科学 2025-11-05 Liuhao Lin , Ke Li , Zihan Xu , Yuchen Shi , Yulei Qin , Yan Zhang , Xing Sun , Rongrong Ji
‹ 上一页 1 8 9 10 下一页 ›