English
Related papers

Related papers: Dynamic Multi-Target Fusion for Efficient Audio-Vi…

200 papers

Current Audio-Visual Source Separation methods primarily adopt two design strategies. The first strategy involves fusing audio and visual features at the bottleneck layer of the encoder, followed by processing the fused features through the…

Sound · Computer Science 2025-05-01 Yinfeng Yu , Shiyu Sun

Efficient aerial data collection is important in many remote sensing applications. In large-scale monitoring scenarios, deploying a team of unmanned aerial vehicles (UAVs) offers improved spatial coverage and robustness against individual…

Robotics · Computer Science 2023-03-03 Jonas Westheider , Julius Rückin , Marija Popović

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

3D vehicle detection based on multi-modal fusion is an important task of many applications such as autonomous driving. Although significant progress has been made, we still observe two aspects that need to be further improvement: First, the…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Zehan Zhang , Ming Zhang , Zhidong Liang , Xian Zhao , Ming Yang , Wenming Tan , ShiLiang Pu

Existing deep learning methods for action recognition in videos require a large number of labeled videos for training, which is labor-intensive and time-consuming. For the same action, the knowledge learned from different media types, e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-19 Yang Liu , Zhaoyang Lu , Jing Li , Tao Yang , Chao Yao

Multimodal camera-LiDAR fusion technology has found extensive application in 3D object detection, demonstrating encouraging performance. However, existing methods exhibit significant performance degradation in challenging scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Sixian Liu , Chen Xu , Qiang Wang , Donghai Shi , Yiwen Li

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Wentao Zhu

With the emergence of varied visual navigation tasks (e.g, image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Hanqing Wang , Wei Liang , Luc Van Gool , Wenguan Wang

Robust navigation in diverse environments and domains requires both accurate state estimation and transparent decision making. We present PhysNav-DG, a novel framework that integrates classical sensor fusion with the semantic power of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Trisanth Srinivasan , Santosh Patapati

Diffusion policies are becoming mainstream in robotic manipulation but suffer from hard negative class imbalance due to uniform sampling and lack of sample difficulty awareness, leading to slow training convergence and frequent inference…

Robotics · Computer Science 2026-04-20 Xinglei Yu , Zhenyang Liu , Shufeng Nan , Simo Wu , Yanwei Fu

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Seongah Kim , Dinh Phu Tran , Hyeontaek Hwang , Saad Wazir , Duc Do Minh , Daeyoung Kim

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing labeled training data quality and cross-modal fusion…

Multimedia · Computer Science 2025-10-03 Yu Sun , Yin Li , Ruixiao Sun , Chunhui Liu , Fangming Zhou , Ze Jin , Linjie Wang , Xiang Shen , Zhuolin Hao , Hongyu Xiong

This report details the methods of the winning entry of the AVDN Challenge in ICCV CLVL 2023. The competition addresses the Aerial Navigation from Dialog History (ANDH) task, which requires a drone agent to associate dialog history with…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Yifei Su , Dong An , Yuan Xu , Kehan Chen , Yan Huang

Image fusion aims to blend complementary information from multiple sensing modalities, yet existing approaches remain limited in robustness, adaptability, and controllability. Most current fusion networks are tailored to specific tasks and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jiayang Li , Chengjie Jiang , Junjun Jiang , Pengwei Liang , Jiayi Ma , Liqiang Nie

The automatic identification system (AIS) and video cameras have been widely exploited for vessel traffic surveillance in inland waterways. The AIS data could provide the vessel identity and dynamic information on vessel position and…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Yu Guo , Ryan Wen Liu , Jingxiang Qu , Yuxu Lu , Fenghua Zhu , Yisheng Lv

Inspired by the fact that humans use diverse sensory organs to perceive the world, sensors with different modalities are deployed in end-to-end driving to obtain the global context of the 3D scene. In previous works, camera and LiDAR inputs…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Qingwen Zhang , Mingkai Tang , Ruoyu Geng , Feiyi Chen , Ren Xin , Lujia Wang

Visual question answering and visual dialogue tasks have been increasingly studied in the multimodal field towards more practical real-world scenarios. A more challenging task, audio visual scene-aware dialogue (AVSD), is proposed to…

Computation and Language · Computer Science 2019-08-15 Yi-Ting Yeh , Tzu-Chuan Lin , Hsiao-Hua Cheng , Yu-Hsuan Deng , Shang-Yu Su , Yun-Nung Chen

Interaction and navigation defined by natural language instructions in dynamic environments pose significant challenges for neural agents. This paper focuses on addressing two challenges: handling long sequence of subtasks, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-26 Alexander Pashevich , Cordelia Schmid , Chen Sun

The extensive application of unmanned aerial vehicles (UAVs) in military reconnaissance, environmental monitoring, and related domains has created an urgent need for accurate and efficient multi-object tracking (MOT) technologies, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Juanqin Liu , Leonardo Plotegher , Eloy Roura , Shaoming He