English
Related papers

Related papers: GPT4Scene: Understand 3D Scenes from Videos with V…

200 papers

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Zhengyuan Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Fan-Yun Sun , Weiyu Liu , Siyi Gu , Dylan Lim , Goutam Bhat , Federico Tombari , Manling Li , Nick Haber , Jiajun Wu

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generalization remains a persistent challenge, particularly when…

Robotics · Computer Science 2025-02-27 Zerui Li , Gengze Zhou , Haodong Hong , Yanyan Shao , Wenqi Lyu , Yanyuan Qiao , Qi Wu

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Amirmohammad Izadi , Mohammad Ali Banayeeanzade , Fatemeh Askari , Ali Rahimiakbar , Mohammad Mahdi Vahedi , Hosein Hasani , Mahdieh Soleymani Baghshah

The ability to construct mental models of the world is a central aspect of understanding. Similarly, visual understanding can be viewed as the ability to construct a representative model of the system depicted in an image. This work…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Sagi Eppel

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wanyue Zhang , Wenxiang Wu , Wang Xu , Jiaxin Luo , Helu Zhi , Yibin Huang , Shuo Ren , Zitao Liu , Jiajun Zhang

Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to…

Computation and Language · Computer Science 2024-02-01 Chenhui Zhang , Sherrie Wang

With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embodied agents that can operate in diverse environments given…

Robotics · Computer Science 2024-11-07 Haochen Zhang , Nader Zantout , Pujith Kachana , Zongyuan Wu , Ji Zhang , Wenshan Wang

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as counting or relational reasoning. To understand the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Zihan Weng , Lucas Gomez , Taylor Whittington Webb , Pouya Bashivan

Addressing the gap in understanding visual comprehension in Large Language Models (LLMs), we designed a challenge-response study, subjecting Google Bard and GPT-Vision to 64 visual tasks, spanning categories like "Visual Situational…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 David Noever , Samantha Elizabeth Miller Noever

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xiaoyan Wang , Zeju Li , Yifan Xu , Jiaxing Qi , Zhifei Yang , Ruifei Ma , Xiangde Liu , Chao Zhang

Scene understanding is critical for various downstream tasks in autonomous driving, including facilitating driver-agent communication and enhancing human-centered explainability of autonomous vehicle (AV) decisions. This paper evaluates the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Mohammed Elhenawy , Shadi Jaradat , Taqwa I. Alhadidi , Huthaifa I. Ashqar , Ahmed Jaber , Andry Rakotonirainy , Mohammad Abu Tami

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of…

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Ye Tian , Yue Yang , Kaixin Ma , Xiaoman Pan , Yangqiu Song , Dong Yu

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yilun Chen , Shuai Yang , Haifeng Huang , Tai Wang , Runsen Xu , Ruiyuan Lyu , Dahua Lin , Jiangmiao Pang

Vision language models (VLMs) are an exciting emerging class of language models (LMs) that have merged classic LM capabilities with those of image processing systems. However, the ways that these capabilities combine are not always…

Computation and Language · Computer Science 2024-07-03 Qiucheng Wu , Handong Zhao , Michael Saxon , Trung Bui , William Yang Wang , Yang Zhang , Shiyu Chang

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation backbones:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Haozhan Shen , Tiancheng Zhao , Kangjia Zhao , Jianwei Yin

Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-of-the-art VLMs still struggle to understand 3D space and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Shuai Wang , Daoan Zhang , Tianyi Bai , Shitong Shao , Jiebo Luo , Jiaheng Wei

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Xinlu Zhang , Yujie Lu , Weizhi Wang , An Yan , Jun Yan , Lianke Qin , Heng Wang , Xifeng Yan , William Yang Wang , Linda Ruth Petzold
‹ Prev 1 4 5 6 7 8 10 Next ›