English
Related papers

Related papers: KEPT: Knowledge-Enhanced Prediction of Trajectorie…

200 papers

Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as promising solutions.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yongdong Luo , Xiawu Zheng , Guilin Li , Shukang Yin , Haojia Lin , Chaoyou Fu , Jinfa Huang , Jiayi Ji , Fei Chao , Jiebo Luo , Rongrong Ji

Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby leveraging complementary clinical information. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Haozhe Luo , Shelley Zixin Shu , Ziyu Zhou , Robert Berke , Mauricio Reyes

With the rapid development of machine learning, autonomous driving has become a hot issue, making urgent demands for more intelligent perception and planning systems. Self-driving cars can avoid traffic crashes with precisely predicted…

Robotics · Computer Science 2021-11-01 Jianbang Liu , Xinyu Mao , Yuqi Fang , Delong Zhu , Max Q. -H. Meng

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Autonomous driving requires generating safe and reliable trajectories from complex multimodal inputs. Traditional modular pipelines separate perception, prediction, and planning, while recent end-to-end (E2E) systems learn them jointly.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Qihang Peng , Xuesong Chen , Chenye Yang , Shaoshuai Shi , Hongsheng Li

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Ashwin Vinod , Shrey Pandit , Aditya Vavre , Linshen Liu

Autonomous driving requires reliable reasoning over fine-grained 3D scene facts. Fine-grained question answering over multi-modal driving observations provides a natural way to evaluate this capability, yet existing perception pipelines and…

Artificial Intelligence · Computer Science 2026-03-24 Ye Tian , Jingyi Zhang , Zihao Wang , Xiaoyuan Ren , Xiaofan Yu , Onat Gungor , Tajana Rosing

End-to-end autonomous driving demonstrates strong planning capabilities with large-scale data but still struggles in complex, rare scenarios due to limited commonsense. In contrast, Large Vision-Language Models (LVLMs) excel in scene…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Bo Jiang , Shaoyu Chen , Bencheng Liao , Xingyu Zhang , Wei Yin , Qian Zhang , Chang Huang , Wenyu Liu , Xinggang Wang

While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Chengen Xie , Chonghao Sima , Tianyu Li , Bin Sun , Junjie Wu , Zhihui Hao , Hongyang Li

Autonomous vehicles demand detailed maps to maneuver reliably through traffic, which need to be kept up-to-date to ensure a safe operation. A promising way to adapt the maps to the ever-changing road-network is to use crowd-sourced data…

Robotics · Computer Science 2024-10-11 Markus Herb , Nassir Navab , Federico Tombari

Autonomous driving requires accurate scene understanding, including road geometry, traffic agents, and their semantic relationships. In online HD map generation scenarios, raster-based representations are well-suited to vision models but…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Zhigang Sun , Yiru Wang , Anqing Jiang , Shuo Wang , Yu Gao , Yuwen Heng , Shouyi Zhang , An He , Hao Jiang , Jinhao Chai , Zichong Gu , Wang Jijun , Shichen Tang , Lavdim Halilaj , Juergen Luettin , Hao Sun

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jingyu Li , Junjie Wu , Dongnan Hu , Xiangkai Huang , Bin Sun , Zhihui Hao , Xianpeng Lang , Xiatian Zhu , Li Zhang

Multimodal large language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-textual data, including images and videos. This study seeks to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Zhenhua Xu , Yujia Zhang , Enze Xie , Zhen Zhao , Yong Guo , Kwan-Yee. K. Wong , Zhenguo Li , Hengshuang Zhao

Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter novel environments. To address this, we propose the Continual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Seongjun Jeong , Gi-Cheon Kang , Seongho Choi , Joochan Kim , Byoung-Tak Zhang

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Visual transfer learning for unseen categories presents an active research topic yet a challenging task, due to the inherent conflict between preserving category-specific representations and acquiring transferable knowledge. Vision-Language…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Xiao Shi , Yangjun Ou , Zhenzhong Chen

CLIP-based prompt tuning enables pretrained Vision-Language Models (VLMs) to efficiently adapt to downstream tasks. Although existing studies have made significant progress, they pay limited attention to changes in the internal attention…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Haoyang Li , Liang Wang , Siyu Zhou , Jiacheng Sun , Jing Jiang , Chao Wang , Guodong Long , Yan Peng

Interpretable communication is essential for safe and trustworthy autonomous driving, yet current vision-language models (VLMs) often operate under idealized assumptions and struggle to capture user intent in real-world scenarios. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Djamahl Etchegaray , Yuxia Fu , Zi Huang , Yadan Luo

Vision-and-language navigation (VLN) aims to build autonomous visual agents that follow instructions and navigate in real scenes. To remember previously visited locations and actions taken, most approaches to VLN implement memory using…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Shizhe Chen , Pierre-Louis Guhur , Cordelia Schmid , Ivan Laptev

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Hao Jiang , Chuan Hu , Yukang Shi , Yuan He , Ke Wang , Xi Zhang , Zhipeng Zhang