English
Related papers

Related papers: Spatial and Visual Perspective-Taking via View Rot…

200 papers

Language grounding aims at linking the symbolic representation of language (e.g., words) into the rich perceptual knowledge of the outside world. The general approach is to embed both textual and visual information into a common space -the…

Computation and Language · Computer Science 2021-09-15 Hassan Shahmohammadi , Hendrik P. A. Lensch , R. Harald Baayen

Visual attention, which assigns weights to image regions according to their relevance to a question, is considered as an indispensable part by most Visual Question Answering models. Although the questions may involve complex relations among…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Chen Zhu , Yanpeng Zhao , Shuaiyi Huang , Kewei Tu , Yi Ma

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Phillip Y. Lee , Jihyeon Je , Chanho Park , Mikaela Angelina Uy , Leonidas Guibas , Minhyuk Sung

Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types of information from visual and textual domain, such as…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Xihui Liu , Zihao Wang , Jing Shao , Xiaogang Wang , Hongsheng Li

Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remote sensing imagery (RSI) requiring exhaustive visual scanning, models tend to rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Gaozhi Zhou , Hu He , Peng Shen , Jipeng Zhang , Liujue Zhang , Linrui Xu , Zeyuan Wang , Ziyu Li , Xuezhi Cui , Wang Guo , Haifeng Li

Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike…

Robotics · Computer Science 2025-01-15 Ying Zheng , Lei Yao , Yuejiao Su , Yi Zhang , Yi Wang , Sicheng Zhao , Yiyi Zhang , Lap-Pui Chau

While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added complexity by requiring models to direct visual attention,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Gabriel Sarch , Snigdha Saha , Naitik Khandelwal , Ayush Jain , Michael J. Tarr , Aviral Kumar , Katerina Fragkiadaki

Visually impaired people could benefit from Visual Question Answering (VQA) systems to interpret text in their surroundings. However, current models often struggle with recognizing text in the photos taken by this population. Through…

Computation and Language · Computer Science 2025-06-05 Hernán Maina , Guido Ivetta , Mateo Lione Stuto , Julian Martin Eisenschlos , Jorge Sánchez , Luciana Benotti

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models…

Robotics · Computer Science 2026-04-06 Jiashu Yang , Yifan Han , Yucheng Xie , Ning Guo , Wenzhao Lian

Vision-language Models (VLMs) have shown remarkable capabilities in advancing general artificial intelligence, yet the irrational encoding of visual positions persists in inhibiting the models' comprehensive perception performance across…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Zhanpeng Chen , Mingxiao Li , Ziyang Chen , Nan Du , Xiaolong Li , Yuexian Zou

In order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects. We propose…

Computer Vision and Pattern Recognition · Computer Science 2019-10-11 Linjie Li , Zhe Gan , Yu Cheng , Jingjing Liu

While users could embody virtual avatars that mirror their physical movements in Virtual Reality, these avatars' motions can be redirected to enable novel interactions. Excessive redirection, however, could break the user's sense of…

Human-Computer Interaction · Computer Science 2025-02-17 Zhipeng Li , Yishu Ji , Ruijia Chen , Tianqi Liu , Yuntao Wang , Yuanchun Shi , Yukang Yan

Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances. Yet existing solutions generally collapse these factors…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yifan Li , Yingda Yin , Lingting Zhu , Weikai Chen , Shengju Qian , Xin Wang , Yanwei Fu

Spatial misalignment caused by variations in poses and viewpoints is one of the most critical issues that hinders the performance improvement in existing person re-identification (Re-ID) algorithms. To address this problem, in this paper,…

Computer Vision and Pattern Recognition · Computer Science 2018-05-17 Qin Zhou , Heng Fan , Hua Yang , Hang Su , Shibao Zheng , Shuang Wu , Haibin Ling

Traditional visual servoing methods suffer from serving between scenes from multiple perspectives, which humans can complete with visual signals alone. In this paper, we investigated how multi-perspective visual servoing could be solved…

Robotics · Computer Science 2023-12-27 Lei Zhang , Jiacheng Pei , Kaixin Bai , Zhaopeng Chen , Jianwei Zhang

Predictive models have been at the core of many robotic systems, from quadrotors to walking robots. However, it has been challenging to develop and apply such models to practical robotic manipulation due to high-dimensional sensory…

Robotics · Computer Science 2020-09-14 Lucas Manuelli , Yunzhu Li , Pete Florence , Russ Tedrake

Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizon multimodal scenarios, often suffering from visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Chenghao Li , Fusheng Hao , Xikai Zhang , Likang Xiao , Yanwei Ren , Fuxiang Wu , Quan Chen , Liu Liu

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e.,…

Robotics · Computer Science 2025-12-23 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2017-02-28 Hanwang Zhang , Zawlin Kyaw , Shih-Fu Chang , Tat-Seng Chua

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Peiyao Wang , Weixin Luo , Yanyu Xu , Haojie Li , Shugong Xu , Jianyu Yang , Shenghua Gao