English
Related papers

Related papers: MTA: Multimodal Task Alignment for BEV Perception …

200 papers

Bird's-Eye-View (BEV) semantic maps have become an essential component of automated driving pipelines due to the rich representation they provide for decision-making tasks. However, existing approaches for generating these maps still follow…

Computer Vision and Pattern Recognition · Computer Science 2023-02-09 Nikhil Gosala , Kürsat Petek , Paulo L. J. Drews-Jr , Wolfram Burgard , Abhinav Valada

Reliable 3D object detection is fundamental to autonomous driving, and multimodal fusion algorithms using cameras and LiDAR remain a persistent challenge. Cameras provide dense visual cues but ill posed depth; LiDAR provides a precise 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Venkatraman Narayanan , Bala Sai , Rahul Ahuja , Pratik Likhar , Varun Ravi Kumar , Senthil Yogamani

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Jinyu Yang , Jiali Duan , Son Tran , Yi Xu , Sampath Chanda , Liqun Chen , Belinda Zeng , Trishul Chilimbi , Junzhou Huang

We present an end-to-end method for object detection and trajectory prediction utilizing multi-view representations of LiDAR returns and camera images. In this work, we recognize the strengths and weaknesses of different view…

Computer Vision and Pattern Recognition · Computer Science 2021-10-20 Sudeep Fadadu , Shreyash Pandey , Darshan Hegde , Yi Shi , Fang-Chieh Chou , Nemanja Djuric , Carlos Vallespi-Gonzalez

This paper introduces BEV-VLM, a novel approach for trajectory planning in autonomous driving that leverages Vision-Language Models (VLMs) with Bird's-Eye View (BEV) feature maps as visual input. Unlike conventional trajectory planning…

Robotics · Computer Science 2026-03-02 Guancheng Chen , Sheng Yang , Tong Zhan , Jian Wang

Public distrust of self-driving cars is growing. Studies emphasize the need for interpreting the behavior of these vehicles to passengers to promote trust in autonomous systems. Interpreters can enhance trust by improving transparency and…

Human-Computer Interaction · Computer Science 2025-01-14 Xuewen Luo , Fan Ding , Ruiqi Chen , Rishikesh Panda , Junnyong Loo , Shuyun Zhang

Autonomous driving requires an accurate representation of the environment. A strategy toward high accuracy is to fuse data from several sensors. Learned Bird's-Eye View (BEV) encoders can achieve this by mapping data from individual sensors…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Thomas Monninger , Vandana Dokkadi , Md Zafar Anwar , Steffen Staab

A robust awareness of how dynamic scenes evolve is essential for Autonomous Driving systems, as they must accurately detect, track, and predict the behaviour of surrounding obstacles. Traditional perception pipelines that rely on modular…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Miguel Antunes-García , Santiago Montiel-Marín , Fabio Sánchez-García , Rodrigo Gutiérrez-Moreno , Rafael Barea , Luis M. Bergasa

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Arijit Ray , Jiafei Duan , Ellis Brown , Reuben Tan , Dina Bashkirova , Rose Hendrix , Kiana Ehsani , Aniruddha Kembhavi , Bryan A. Plummer , Ranjay Krishna , Kuo-Hao Zeng , Kate Saenko

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have…

Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent self-attention based methods struggle to effectively handle cross-modal semantic alignments between vision and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Qiangguo Jin , Xianyao Zheng , Hui Cui , Changming Sun , Yuqi Fang , Cong Cong , Ran Su , Leyi Wei , Ping Xuan , Junbo Wang

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jianhua Han , Meng Tian , Jiangtong Zhu , Fan He , Huixin Zhang , Sitong Guo , Dechang Zhu , Hao Tang , Pei Xu , Yuze Guo , Minzhe Niu , Haojie Zhu , Qichao Dong , Xuechao Yan , Siyuan Dong , Lu Hou , Qingqiu Huang , Xiaosong Jia , Hang Xu

Accurate localization ability is fundamental in autonomous driving. Traditional visual localization frameworks approach the semantic map-matching problem with geometric models, which rely on complex parameter tuning and thus hinder…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Zhihuang Zhang , Meng Xu , Wenqiang Zhou , Tao Peng , Liang Li , Stefan Poslad

Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation critical for downstream planning and control. However, real-world deployment faces challenges…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Lifeng Zhuo , Kefan Jin , Zhe Liu , Hesheng Wang

We present VASTA, a novel vision and language-assisted Programming By Demonstration (PBD) system for smartphone task automation. Development of a robust PBD automation system requires overcoming three key challenges: first, how to make a…

Human-Computer Interaction · Computer Science 2019-11-06 Alborz Rezazadeh Sereshkeh , Gary Leung , Krish Perumal , Caleb Phillips , Minfan Zhang , Afsaneh Fazly , Iqbal Mohomed

The perception system for autonomous driving generally requires to handle multiple diverse sub-tasks. However, current algorithms typically tackle individual sub-tasks separately, which leads to low efficiency when aiming at obtaining…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xuesong Chen , Shaoshuai Shi , Tao Ma , Jingqiu Zhou , Simon See , Ka Chun Cheung , Hongsheng Li

Semantic Bird's Eye View (BEV) maps offer a rich representation with strong occlusion reasoning for various decision making tasks in autonomous driving. However, most BEV mapping approaches employ a fully supervised learning paradigm that…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Nikhil Gosala , Kürsat Petek , B Ravi Kiran , Senthil Yogamani , Paulo Drews-Jr , Wolfram Burgard , Abhinav Valada

Visual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal and imbalanced overall performance due to latent knowledge…

Artificial Intelligence · Computer Science 2026-01-22 Yanqi Dai , Yong Wang , Zebin You , Dong Jing , Xiangxiang Chu , Zhiwu Lu

Recently, video captioning has been attracting an increasing amount of interest, due to its potential for improving accessibility and information retrieval. While existing methods rely on different kinds of visual features and model…

Computer Vision and Pattern Recognition · Computer Science 2016-12-02 Xiang Long , Chuang Gan , Gerard de Melo

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modalities at different…

Computation and Language · Computer Science 2018-04-17 Xin Wang , Yuan-Fang Wang , William Yang Wang
‹ Prev 1 8 9 10 Next ›