English
Related papers

Related papers: Multi-level Attention Fusion Network for Audio-vis…

200 papers

Image classification models often demonstrate unstable performance in real-world applications due to variations in image information, driven by differing visual perspectives of subject objects and lighting discrepancies. To mitigate these…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Yuze Zheng , Zixuan Li , Xiangxian Li , Jinxing Liu , Yuqing Wang , Xiangxu Meng , Lei Meng

The increasing use of synthetic media, particularly deepfakes, is an emerging challenge for digital content verification. Although recent studies use both audio and visual information, most integrate these cues within a single model, which…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Sayeem Been Zaman , Wasimul Karim , Arefin Ittesafun Abian , Reem E. Mohamed , Md Rafiqul Islam , Asif Karim , Sami Azam

Multi-source data classification is a critical yet challenging task for remote sensing image interpretation. Existing methods lack adaptability to diverse land cover types when modeling frequency domain features. To this end, we propose a…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Yikang Zhao , Feng Gao , Xuepeng Jin , Junyu Dong , Qian Du

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Ziheng Zhou , Jinxing Zhou , Wei Qian , Shengeng Tang , Xiaojun Chang , Dan Guo

Emotion Recognition in Conversations (ERC) has considerable prospects for developing empathetic machines. For multimodal ERC, it is vital to understand context and fuse modality information in conversations. Recent graph-based fusion…

Computation and Language · Computer Science 2022-03-07 Dou Hu , Xiaolong Hou , Lingwei Wei , Lianxin Jiang , Yang Mo

Predicting saliency in videos is a challenging problem due to complex modeling of interactions between spatial and temporal information, especially when ever-changing, dynamic nature of videos is considered. Recently, researchers have…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Aysun Kocak , Erkut Erdem , Aykut Erdem

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lacking cross-modal…

Multimedia · Computer Science 2025-10-16 Huilai Li , Yonghao Dang , Ying Xing , Yiming Wang , Jianqin Yin

Multi-modality data is becoming readily available in remote sensing (RS) and can provide complementary information about the Earth's surface. Effective fusion of multi-modal information is thus important for various applications in RS, but…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Qinghui Liu , Michael Kampffmeyer , Robert Jenssen , Arnt-Børre Salberg

This paper strives for video event detection using a representation learned from deep convolutional neural networks. Different from the leading approaches, who all learn from the 1,000 classes defined in the ImageNet Large Scale Visual…

Computer Vision and Pattern Recognition · Computer Science 2017-12-14 Pascal Mettes , Dennis C. Koelma , Cees G. M. Snoek

Depression has been the leading cause of mental-health illness worldwide. Major depressive disorder (MDD), is a common mental health disorder that affects both psychologically as well as physically which could lead to loss of lives. Due to…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Anupama Ray , Siddharth Kumar , Rutvik Reddy , Prerana Mukherjee , Ritu Garg

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

In this paper, the dual-optical attention fusion crowd head point counting model (TAPNet) is proposed to address the problem of the difficulty of accurate counting in complex scenes such as crowd dense occlusion and low light in crowd…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Fei Zhou , Yi Li , Mingqing Zhu

In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data driven techniques, e.g., deep learning, where audio…

Sound · Computer Science 2018-06-29 Eduardo Fonseca , Rong Gong , Xavier Serra

The weakly supervised sound event detection problem is the task of predicting the presence of sound events and their corresponding starting and ending points in a weakly labeled dataset. A weak dataset associates each training sample (a…

Sound · Computer Science 2021-06-22 Mohammad Rasool Izadi , Robert Stevenson , Laura N. Kloepper

In recent years, there have been numerous developments towards solving multimodal tasks, aiming to learn a stronger representation than through a single modality. Certain aspects of the data can be particularly useful in this case - for…

Machine Learning · Statistics 2023-09-06 Cătălina Cangea , Petar Veličković , Pietro Liò

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

This paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Siqi Li , Changqing Zou , Yipeng Li , Xibin Zhao , Yue Gao

Unmanned aerial vehicles (UAVs) are now widely applied to data acquisition due to its low cost and fast mobility. With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this,…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Pu Jin , Lichao Mou , Yuansheng Hua , Gui-Song Xia , Xiao Xiang Zhu

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

In video compression, most of the existing deep learning approaches concentrate on the visual quality of a single frame, while ignoring the useful priors as well as the temporal information of adjacent frames. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2019-01-16 Xiandong Meng , Xuan Deng , Shuyuan Zhu , Shuaicheng Liu , Chuan Wang , Chen Chen , Bing Zeng