中文
相关论文

相关论文: SpaRRTa: A Synthetic Benchmark for Evaluating Spat…

200 篇论文

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Yifan Liu , Fangneng Zhan , Kaichen Zhou , Yilun Du , Paul Pu Liang , Hanspeter Pfister

State Space Models (SSMs) have recently emerged as an alternative to Vision Transformers (ViTs) due to their unique ability of modeling global relationships with linear complexity. SSMs are specifically designed to capture spatially…

Foundation Models (FMs), e.g., large language models, possess attributes of intelligence which offer promise to endow a robot with the contextual understanding necessary to navigate complex, unstructured tasks in the wild. We see three core…

The rapid advancements in Vision-Language Models (VLMs) have shown great potential in tackling mathematical reasoning tasks that involve visual context. Unlike humans who can reliably apply solution steps to similar problems with minor…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Chengke Zou , Xingang Guo , Rui Yang , Junyu Zhang , Bin Hu , Huan Zhang

Although visual foundation models like DINOv2 provide state-of-the-art performance as feature extractors, their complex, high-dimensional representations create substantial hurdles for interpretability. This work proposes DINO-QPM, which…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Robert Zimmermann , Thomas Norrenbrock , Bodo Rosenhahn

Spatial reasoning is a crucial component of both biological and artificial intelligence. In this work, we present a comprehensive study of the capability of current state-of-the-art large language models (LLMs) on spatial reasoning. To…

计算与语言 · 计算机科学 2024-06-10 Md Imbesat Hassan Rizvi , Xiaodan Zhu , Iryna Gurevych

Spatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of…

计算机视觉与模式识别 · 计算机科学 2020-09-03 Wenyu Han , Siyuan Xiang , Chenhui Liu , Ruoyu Wang , Chen Feng

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Sashuai Zhou , Qiang Zhou , Junpeng Ma , Yue Cao , Ruofan Hu , Ziang Zhang , Xiaoda Yang , Zhibin Wang , Jun Song , Cheng Yu , Bo Zheng , Zhou Zhao

Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language…

计算与语言 · 计算机科学 2025-04-18 Zheyuan Zhang , Fengyuan Hu , Jayjun Lee , Freda Shi , Parisa Kordjamshidi , Joyce Chai , Ziqiao Ma

CAPTCHA, originally designed to distinguish humans from robots, has evolved into a real-world benchmark for assessing the spatial reasoning capabilities of vision-language models. In this work, we first show that step-by-step reasoning is…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Python Song , Luke Tenyi Chang , Yun-Yun Tsai , Penghui Li , Junfeng Yang

Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear. In this paper, we benchmark popular MFMs (GPT-4o, o4-mini,…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Rahul Ramachandran , Ali Garjani , Roman Bachmann , Andrei Atanov , Oğuzhan Fatih Kar , Amir Zamir

Vision Language Models (VLMs) excel at identifying and describing objects but often fail at spatial reasoning. We study why VLMs, such as LLaVA, underutilize spatial cues despite having positional encodings and spatially rich vision encoder…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Jianing Qi , Jiawei Liu , Hao Tang , Zhigang Zhu

Structure-from-Motion (SfM), a task aiming at jointly recovering camera poses and 3D geometry of a scene given a set of images, remains a hard problem with still many open challenges despite decades of significant progress. The traditional…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Bardienus Duisterhof , Lojze Zust , Philippe Weinzaepfel , Vincent Leroy , Yohann Cabon , Jerome Revaud

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Junfei Wu , Jian Guan , Kaituo Feng , Qiang Liu , Shu Wu , Liang Wang , Wei Wu , Tieniu Tan

Robotic task planning in real-world environments requires not only object recognition but also a nuanced understanding of spatial relationships between objects. We present a spatial-relationship-aware dataset of nearly 1,000 robot-acquired…

机器人学 · 计算机科学 2025-06-17 Peng Wang , Minh Huy Pham , Zhihao Guo , Wei Zhou

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Mingrui Wu , Zhaozhi Wang , Fangjinhua Wang , Jiaolong Yang , Marc Pollefeys , Tong Zhang

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Kexin Tian , Jingrui Mao , Yunlong Zhang , Jiwan Jiang , Yang Zhou , Zhengzhong Tu

Few-shot semantic segmentation (FSS) is a crucial challenge in computer vision, driving extensive research into a diverse range of methods, from advanced meta-learning techniques to simple transfer learning baselines. With the emergence of…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Reda Bensaid , Vincent Gripon , François Leduc-Primeau , Lukas Mauch , Ghouthi Boukli Hacene , Fabien Cardinaux

Vision Foundation Models (VFMs) and Vision Language Models (VLMs) have revolutionized computer vision by providing rich semantic and geometric representations. This paper presents a comprehensive visual comparison between CLIP based and…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Md Selim Sarowar , Sungho Kim

Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tuning incurs prohibitive computational and storage overhead. Parameter-Efficient…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Lingyu Xiong , Jinjin Shi , Xuran Xu , Cong Luo , Runyu Shi , Ying Huang
‹ 上一页 1 8 9 10 下一页 ›