English
Related papers

Related papers: PAC Bench: Do Foundation Models Understand Prerequ…

200 papers

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Phillip Y. Lee , Jihyeon Je , Chanho Park , Mikaela Angelina Uy , Leonidas Guibas , Minhyuk Sung

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive…

Robotics · Computer Science 2025-12-23 Max Argus , Jelena Bratulic , Houman Masnavi , Maxim Velikanov , Nick Heppert , Abhinav Valada , Thomas Brox

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel benchmark to assess…

Robotics · Computer Science 2025-03-04 Ricardo Garcia , Shizhe Chen , Cordelia Schmid

Vision-and-Language Navigation (VLN) refers to the task of enabling autonomous robots to navigate unfamiliar environments by following natural language instructions. While recent Large Vision-Language Models (LVLMs) have shown promise in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Vebjørn Haug Kåsene , Pierre Lison

Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Parker Liu , Chenxin Li , Zhengxin Li , Yipeng Wu , Wuyang Li , Zhiqin Yang , Zhenyuan Zhang , Yunlong Lin , Sirui Han , Brandon Y. Feng

Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities beyond narrow examination success. Current benchmarks,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Dingyu Wang , Zimu Yuan , Jiajun Liu , Shanggui Liu , Nan Zhou , Tianxing Xu , Di Huang , Dong Jiang

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR…

Humans naturally possess the spatial reasoning ability to form and manipulate images and structures of objects in space. There is an increasing effort to endow Vision-Language Models (VLMs) with similar spatial reasoning capabilities.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jiahuan Zhang , Shunwen Bai , Tianheng Wang , Kaiwen Guo , Kai Han , Guozheng Rao , Kaicheng Yu

Recent advances in microscopy have enabled the rapid generation of terabytes of image data in cell biology and biomedical research. Vision-language models (VLMs) offer a promising solution for large-scale biological image analysis,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Alejandro Lozano , Jeffrey Nirschl , James Burgess , Sanket Rajan Gupte , Yuhui Zhang , Alyssa Unell , Serena Yeung-Levy

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Dhruba Ghosh , Yuhui Zhang , Ludwig Schmidt

Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-modal data. To compensate for the deficiency of robot data,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Dejie Yang , Zijing Zhao , Yang Liu

Vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However, understanding fine-grained visual-linguistic concepts, such as attributes and inter-object relationships, remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Wujian Peng , Sicheng Xie , Zuyao You , Shiyi Lan , Zuxuan Wu

Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment. In this work, we systematically evaluate single-view object captioning for tabletop scenes…

Robotics · Computer Science 2026-04-24 Federico Tavella , Amber Drinkwater , Angelo Cangelosi

Vision-Language Models (VLMs) have emerged as general purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, also lacking some basic visual…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Shivam Chandhok , Wan-Cyuan Fan , Leonid Sigal

Vision Language Models (VLMs) have received significant attention in recent years in the robotics community. VLMs are shown to be able to perform complex visual reasoning and scene understanding tasks, which makes them regarded as a…

Robotics · Computer Science 2024-06-14 Siyuan Huang , Haonan Chang , Yuhan Liu , Yimeng Zhu , Hao Dong , Peng Gao , Abdeslam Boularias , Hongsheng Li

In recent years, vision-language models (VLMs) have shown remarkable performance on visual reasoning tasks (e.g. attributes, location). While such tasks measure the requisite knowledge to ground and reason over a given visual instance, they…

Computation and Language · Computer Science 2022-09-16 Shikhar Singh , Ehsan Qasemi , Muhao Chen

Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Baiqi Li , Zhiqiu Lin , Wenxuan Peng , Jean de Dieu Nyandwi , Daniel Jiang , Zixian Ma , Simran Khanuja , Ranjay Krishna , Graham Neubig , Deva Ramanan

Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene through multiple…

Robotics · Computer Science 2025-12-23 Jin Wang , Kim Tien Ly , Jacques Cloete , Nikos Tsagarakis , Ioannis Havoutis

A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation tasks typically assume that instructions are perfectly aligned…

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to…

‹ Prev 1 4 5 6 7 8 10 Next ›