中文
相关论文

相关论文: SPHERE: Unveiling Spatial Blind Spots in Vision-La…

200 篇论文

The multimedia community has shown a significant interest in perceiving and representing the physical world with multimodal pretrained neural network models, and among them, the visual-language pertaining (VLP) is, currently, the most…

多媒体 · 计算机科学 2023-08-28 Fei Wang , Liang Ding , Jun Rao , Ye Liu , Li Shen , Changxing Ding

Spatial confounding poses a significant challenge in scientific studies involving spatial data, where unobserved spatial variables can influence both treatment and outcome, possibly leading to spurious associations. To address this problem,…

Autonomous robotic systems require spatio-temporal understanding of dynamic environments to ensure reliable navigation and interaction. While Vision-Language Models (VLMs) provide open-world semantic priors, they lack grounding in 3D…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

Spatial reasoning, the ability to ground language in 3D understanding, remains a persistent challenge for Vision-Language Models (VLMs). We identify two fundamental bottlenecks: inadequate 3D understanding capabilities stemming from…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Yejie Guo , Yunzhong Hou , Wufei Ma , Meng Tang , Ming-Hsuan Yang

Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Although these numbers appear meaningful, it remains unclear…

人工智能 · 计算机科学 2026-05-25 Jianshu Zhang , Yijiang Li , Huifeixin Chen , Haoran Lu , Letian Xue , Bingyang Wang , Han Liu

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Wufei Ma , Haoyu Chen , Guofeng Zhang , Yu-Cheng Chou , Jieneng Chen , Celso M de Melo , Alan Yuille

Many new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the…

计算机视觉与模式识别 · 计算机科学 2019-12-19 Jeonghun Baek , Geewook Kim , Junyeop Lee , Sungrae Park , Dongyoon Han , Sangdoo Yun , Seong Joon Oh , Hwalsuk Lee

Large language models (LLMs) show strong performance across natural language processing (NLP), mathematical reasoning, and programming, and recent large reasoning models (LRMs) further emphasize explicit reasoning. Yet their computational…

人工智能 · 计算机科学 2025-10-13 Hyundong Jin , Joonghyuk Hahn , Yo-Sub Han

Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Turhan Can Kargin , Wojciech Jasiński , Adam Pardyl , Bartosz Zieliński , Marcin Przewięźlikowski

Process or step-wise supervision has played a crucial role in advancing complex multi-step reasoning capabilities of Large Language Models (LLMs). However, efficient, high-quality automated process annotation remains a significant…

计算与语言 · 计算机科学 2026-03-03 Md Imbesat Hassan Rizvi , Xiaodan Zhu , Iryna Gurevych

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Ziang Zhang , Zehan Wang , Guanghao Zhang , Weilong Dai , Yan Xia , Ziang Yan , Minjie Hong , Zhou Zhao

Recent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Zhenhua Ning , Zhuotao Tian , Shaoshuai Shi , Guangming Lu , Daojing He , Wenjie Pei , Li Jiang

The rapid advancement of Large Vision-Language Models (VLMs), both general-domain models and those specifically tailored for remote sensing, has demonstrated exceptional perception and reasoning capabilities in Earth observation tasks.…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Xiao An , Jiaxing Sun , Zihan Gui , Wei He

Understanding the spatial relations between objects in images is a surprisingly challenging task. A chair may be "behind" a person even if it appears to the left of the person in the image (depending on which way the person is facing). Two…

计算机视觉与模式识别 · 计算机科学 2019-09-02 Kaiyu Yang , Olga Russakovsky , Jia Deng

Referring image segmentation aims to produce a pixel-level mask for the image region described by a natural-language expression. Although pretrained vision-language models have improved semantic grounding, many existing methods still rely…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Alaa Dalaq , Muzammil Behzad

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained vision language models (VLMs), recent approaches are still not qualified to accurately…

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

In cognitive science and AI, a longstanding question is whether machines learn representations that align with those of the human mind. While current models show promise, it remains an open question whether this alignment is superficial or…

神经元与认知 · 定量生物学 2025-10-27 Craig Sanders , Billy Dickson , Sahaj Singh Maini , Robert Nosofsky , Zoran Tiganj

Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning…

计算与语言 · 计算机科学 2025-10-01 Julius Mayer , Mohamad Ballout , Serwan Jassim , Farbod Nosrat Nezami , Elia Bruni

Large language models (LLMs) require constant updates to remain aligned with evolving real-world knowledge. Model editing offers a lightweight alternative to retraining, but sequential editing often destabilizes representations and induces…

计算与语言 · 计算机科学 2026-05-15 Qingyuan Liu , Jia-Chen Gu , Yunzhi Yao , Hong Wang , Nanyun Peng