English
Related papers

Related papers: MTA: Multimodal Task Alignment for BEV Perception …

200 papers

In the field of 3D object detection tasks, fusing heterogeneous features from LiDAR and camera sensors into a unified Bird's Eye View (BEV) representation is a widely adopted paradigm. However, existing methods often suffer from imprecise…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Ziying Song , Hongyu Pan , Feiyang Jia , Yongchang Zhang , Lin Liu , Lei Yang , Shaoqing Xu , Peiliang Wu , Caiyan Jia , Zheng Zhang , Yadan Luo

Infrastructure-based perception plays a crucial role in intelligent transportation systems, offering global situational awareness and enabling cooperative autonomy. However, existing camera-based detection models often underperform in such…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Yun Zhang , Zhaoliang Zheng , Johnson Liu , Zhiyu Huang , Zewei Zhou , Zonglin Meng , Tianhui Cai , Jiaqi Ma

Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate naturalistic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with visual cues including…

Computation and Language · Computer Science 2025-10-27 Sam O'Connor Russell , Naomi Harte

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Zhengyuan Yang , Yijuan Lu , Jianfeng Wang , Xi Yin , Dinei Florencio , Lijuan Wang , Cha Zhang , Lei Zhang , Jiebo Luo

Recent advancements in bird's eye view (BEV) representations have shown remarkable promise for in-vehicle 3D perception. However, while these methods have achieved impressive results on standard benchmarks, their robustness in varied…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Shaoyuan Xie , Lingdong Kong , Wenwei Zhang , Jiawei Ren , Liang Pan , Kai Chen , Ziwei Liu

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot manipulation. However, due to the heterogeneity of tactile…

Robotics · Computer Science 2025-08-25 Zhengxue Cheng , Yiqian Zhang , Wenkang Zhang , Haoyu Li , Keyu Wang , Li Song , Hengdi Zhang

Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Gyusam Chang , Wonseok Roh , Sujin Jang , Dongwook Lee , Daehyun Ji , Gyeongrok Oh , Jinsun Park , Jinkyu Kim , Sangpil Kim

As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual and textual models,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Yan Ling , Jianfei Yu , Rui Xia

Multimodal Machine Translation (MMT) focuses on enhancing text-only translation with visual features, which has attracted considerable attention from both natural language processing and computer vision communities. Recent advances still…

Computation and Language · Computer Science 2022-11-29 Hongcheng Guo , Jiaheng Liu , Haoyang Huang , Jian Yang , Zhoujun Li , Dongdong Zhang , Zheng Cui , Furu Wei

Video captioning aims to describe video contents using natural language format that involves understanding and interpreting scenes, actions and events that occurs simultaneously on the view. Current approaches have mainly concentrated on…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Antoine Hanna-Asaad , Decky Aspandi , Titus Zaharia

In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Yuexin Ma , Tai Wang , Xuyang Bai , Huitong Yang , Yuenan Hou , Yaming Wang , Yu Qiao , Ruigang Yang , Dinesh Manocha , Xinge Zhu

Multi-view 3D object detection is becoming popular in autonomous driving due to its high effectiveness and low cost. Most of the current state-of-the-art detectors follow the query-based bird's-eye-view (BEV) paradigm, which benefits from…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Zhangyang Qi , Jiaqi Wang , Xiaoyang Wu , Hengshuang Zhao

Multi-task learning has emerged as a powerful paradigm to solve a range of tasks simultaneously with good efficiency in both computation resources and inference time. However, these algorithms are designed for different tasks mostly not…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Xiwen Liang , Minzhe Niu , Jianhua Han , Hang Xu , Chunjing Xu , Xiaodan Liang

In autonomous driving, multi-modal perception tasks like 3D object detection typically rely on well-synchronized sensors, both at training and inference. However, despite the use of hardware- or software-based synchronization algorithms,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Shiming Wang , Holger Caesar , Liangliang Nan , Julian F. P. Kooij

LiDAR and camera are two essential sensors for 3D object detection in autonomous driving. LiDAR provides accurate and reliable 3D geometry information while the camera provides rich texture with color. Despite the increasing popularity of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Qi Jiang , Hao Sun , Xi Zhang

In recent years, multimodal learning has become essential in robotic vision and information fusion, especially for understanding human behavior in complex environments. However, current methods struggle to fully leverage the textual…

Robotics · Computer Science 2025-09-23 Yanxin Zhang , Liang He , Zeyi Kang , Zuheng Ming , Kaixing Zhao

Estimating a semantically segmented bird's-eye-view (BEV) map from a single image has become a popular technique for autonomous control and navigation. However, they show an increase in localization error with distance from the camera.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Avishkar Saha , Oscar Mendez , Chris Russell , Richard Bowden

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Ning Cheng , You Li , Jing Gao , Bin Fang , Jinan Xu , Wenjuan Han

Vision-based bird's-eye-view (BEV) 3D object detection has advanced significantly in autonomous driving by offering cost-effectiveness and rich contextual information. However, existing methods often construct BEV representations by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jicheng Yuan , Manh Nguyen Duc , Qian Liu , Manfred Hauswirth , Danh Le Phuoc