English
Related papers

Related papers: Is a 3D-Tokenized LLM the Key to Reliable Autonomo…

200 papers

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yunze Man , Shihao Wang , Guowen Zhang , Johan Bjorck , Zhiqi Li , Liang-Yan Gui , Jim Fan , Jan Kautz , Yu-Xiong Wang , Zhiding Yu

Recent advances in Multimodal Large Language Models (MLLMs) have expanded reasoning capabilities into 3D domains, enabling fine-grained spatial understanding. However, the substantial size of 3D MLLMs and the high dimensionality of input…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Yuhui Lin , Siyue Yu , Yuxing Yang , Guangliang Cheng , Jimin Xiao

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Vision-Language Models (VLMs) and Multi-Modal Language models (MMLMs) have become prominent in autonomous driving research, as these models can provide interpretable textual reasoning and responses for end-to-end autonomous driving safety…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Akshay Gopalkrishnan , Ross Greer , Mohan Trivedi

Autonomous Driving (AD) encounters significant safety hurdles in long-tail unforeseen driving scenarios, largely stemming from the non-interpretability and poor generalization of the deep neural networks within the AD system, particularly…

Artificial Intelligence · Computer Science 2024-03-25 Yixuan Wang , Ruochen Jiao , Sinong Simon Zhan , Chengtian Lang , Chao Huang , Zhaoran Wang , Zhuoran Yang , Qi Zhu

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensors. However,…

Robotics · Computer Science 2025-06-23 Junwei You , Haotian Shi , Zhuoyu Jiang , Zilin Huang , Rui Gan , Keshu Wu , Xi Cheng , Xiaopeng Li , Bin Ran

With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This convergence enables…

Robotics · Computer Science 2025-11-19 Vinit Mehta , Charu Sharma , Karthick Thiyagarajan

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Peizheng Li , Zhenghao Zhang , David Holtz , Hang Yu , Yutong Yang , Yuzhi Lai , Rui Song , Andreas Geiger , Andreas Zell

Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control systems. However, their success is tied to one core capability, reliable object detection in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Sayed Pedram Haeri Boroujeni , Niloufar Mehrabi , Hazim Alzorgan , Mahlagha Fazeli , Abolfazl Razi

Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Grounding. Unlike 2D VLMs that typically process images through…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Haoyuan Li , Yanpeng Zhou , Yufei Gao , Tao Tang , Jianhua Han , Yujie Yuan , Dave Zhenyu Chen , Jiawang Bian , Hang Xu , Xiaodan Liang

End-to-end autonomous driving systems built on Vision Language Models (VLMs) have shown significant promise, yet their reliance on autoregressive architectures introduces some limitations for real-world applications. The sequential,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Can Cui , Yupeng Zhou , Juntong Peng , Sung-Yeon Park , Zichong Yang , Prashanth Sankaranarayanan , Jiaru Zhang , Ruqi Zhang , Ziran Wang

Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeniable candidates for end-to-end driving systems. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Shuo Xing , Hongyuan Hua , Xiangbo Gao , Shenzhe Zhu , Renjie Li , Kexin Tian , Xiaopeng Li , Heng Huang , Tianbao Yang , Zhangyang Wang , Yang Zhou , Huaxiu Yao , Zhengzhong Tu

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Xiaoyu Tian , Junru Gu , Bailin Li , Yicheng Liu , Yang Wang , Zhiyong Zhao , Kun Zhan , Peng Jia , Xianpeng Lang , Hang Zhao

Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities…

Computation and Language · Computer Science 2025-03-28 Yue Li , Meng Tian , Zhenyu Lin , Jiangtong Zhu , Dechang Zhu , Haiqiang Liu , Zining Wang , Yueyi Zhang , Zhiwei Xiong , Xinhai Zhao

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jianhua Han , Meng Tian , Jiangtong Zhu , Fan He , Huixin Zhang , Sitong Guo , Dechang Zhu , Hao Tang , Pei Xu , Yuze Guo , Minzhe Niu , Haojie Zhu , Qichao Dong , Xuechao Yan , Siyuan Dong , Lu Hou , Qingqiu Huang , Xiaosong Jia , Hang Xu

Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yiwei Zhang , Xuesong Chen , Jin Gao , Hanshi Wang , Fudong Ge , Weiming Hu , Shaoshuai Shi , Zhipeng Zhang

End-to-end autonomous driving has emerged as a promising paradigm integrating perception, decision-making, and control within a unified learning framework. Recently, Vision-Language Models (VLMs) have gained significant attention for their…

Robotics · Computer Science 2026-02-05 Yuxuan Han , Kunyuan Wu , Qianyi Shao , Renxiang Xiao , Zilu Wang , Cansen Jiang , Yi Xiao , Liang Hu , Yunjiang Lou

Autonomous vehicles (AVs) require reliable traffic sign recognition and robust lane detection capabilities to ensure safe navigation in complex and dynamic environments. This paper introduces an integrated approach combining advanced deep…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Chandan Kumar Sah , Ankit Kumar Shaw , Xiaoli Lian , Arsalan Shahid Baig , Tuopu Wen , Kun Jiang , Mengmeng Yang , Diange Yang

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Zero-shot vision-and-language navigation (VLN) has gained significant attention due to its minimal data collection costs and inherent generalization. This paradigm is typically driven by the integration of pre-trained Vision-Language Models…

Robotics · Computer Science 2026-05-15 Ziyi Xia , Chaoran Xiong , Litao Wei , Xinhao Hu , Ling Pei