English
Related papers

Related papers: Structured Labeling Enables Faster Vision-Language…

200 papers

Despite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Mohamed Aghzal , Xiang Yue , Erion Plaku , Ziyu Yao

Autonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform…

Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Haruki Sakajo , Hiroshi Takato , Hiroshi Tsutsui , Komei Soda , Hidetaka Kamigaito , Taro Watanabe

Vision-language models (VLMs) integrate visual and textual information, enabling a wide range of applications such as image captioning and visual question answering, making them crucial for modern AI systems. However, their high…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Gaurav Shinde , Anuradha Ravi , Emon Dey , Shadman Sakib , Milind Rampure , Nirmalya Roy

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply…

Computer Vision and Pattern Recognition · Computer Science 2021-03-08 Hanqing Wang , Wenguan Wang , Wei Liang , Caiming Xiong , Jianbing Shen

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Ye Tian , Yue Yang , Kaixin Ma , Xiaoman Pan , Yangqiu Song , Dong Yu

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez

Recent advances in multi-modal large language models (MLLMs) have demonstrated strong performance across various domains; however, their ability to comprehend driving scenes remains less proven. The complexity of driving scenarios, which…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Sung-Yeon Park , Can Cui , Yunsheng Ma , Ahmadreza Moradipari , Rohit Gupta , Kyungtae Han , Ziran Wang

CAR-Scenes is a frame-level dataset for autonomous driving that enables training and evaluation of vision-language models (VLMs) for interpretable, scene-level understanding. We annotate 5,192 images drawn from Argoverse 1, Cityscapes,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Yuankai He , Weisong Shi

Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeniable candidates for end-to-end driving systems. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Shuo Xing , Hongyuan Hua , Xiangbo Gao , Shenzhe Zhu , Renjie Li , Kexin Tian , Xiaopeng Li , Heng Huang , Tianbao Yang , Zhangyang Wang , Yang Zhou , Huaxiu Yao , Zhengzhong Tu

Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Ruiqi Song , Xianda Guo , Yanlun Peng , Qinggong Wei , Hangbin Wu , Long Chen

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Manyi Yao , Bingbing Zhuang , Sparsh Garg , Amit Roy-Chowdhury , Christian Shelton , Manmohan Chandraker , Abhishek Aich

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in which the image-text pairs comprise short texts. To evaluate…

Computation and Language · Computer Science 2025-10-16 Jesse Atuhurra , Iqra Ali , Tomoya Iwakura , Hidetaka Kamigaito , Tatsuya Hiraoka

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Pei Liu , Haipeng Liu , Haichao Liu , Xin Liu , Jinxin Ni , Jun Ma

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jingyu Li , Junjie Wu , Dongnan Hu , Xiangkai Huang , Bin Sun , Zhihui Hao , Xianpeng Lang , Xiatian Zhu , Li Zhang

Autonomous vehicles (AVs) rely on sophisticated perception systems to interpret their surroundings, a cornerstone for safe navigation and decision-making. The integration of Large Language Models (LLMs) into AV perception frameworks offers…

Robotics · Computer Science 2024-12-31 Athanasios Karagounis