English
Related papers

Related papers: GateFusion: Hierarchical Gated Cross-Modal Fusion …

200 papers

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Fusing LiDAR and camera information is essential for achieving accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Yang Jiao , Zequn Jie , Shaoxiang Chen , Jingjing Chen , Lin Ma , Yu-Gang Jiang

Existing works on weakly-supervised audio-visual video parsing adopt hybrid attention network (HAN) as the multi-modal embedding to capture the cross-modal context. It embeds the audio and visual modalities with a shared network, where the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Yating Xu , Conghui Hu , Gim Hee Lee

The fusion of multimodal sensor data streams such as camera images and lidar point clouds plays an important role in the operation of autonomous vehicles (AVs). Robust perception across a range of adverse weather and lighting conditions is…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Shounak Sural , Nishad Sahu , Ragunathan Rajkumar

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Wentao Zhu

Detecting hateful content in multimodal memes presents unique challenges, as harmful messages often emerge from the complex interplay between benign images and text. We propose GatedCLIP, a Vision-Language model that enhances CLIP's…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Yingying Guo , Ke Zhang , Zirong Zeng

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Vinaya Sree Katamneni , Ajita Rattani

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from…

Sound · Computer Science 2024-09-12 Zehao Wang , Haobo Yue , Zhicheng Zhang , Da Mu , Jin Tang , Jianqin Yin

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

Sound · Computer Science 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

Automatic Speaker Verification (ASV) systems, which identify speakers based on their voice characteristics, have numerous applications, such as user authentication in financial transactions, exclusive access control in smart devices, and…

Automated audio captioning (AAC) which generates textual descriptions of audio content. Existing AAC models achieve good results but only use the high-dimensional representation of the encoder. There is always insufficient information…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Jianyuan Sun , Xubo Liu , Xinhao Mei , Volkan Kılıç , Mark D. Plumbley , Wenwu Wang

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Ina Salaj , Arijit Biswas

State-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input. However, the main hypothesis in this paper is that body dynamics is also highly correlated to "speaking" (and "listening") actions and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Tiago Roxo , Joana C. Costa , Pedro Inácio , Hugo Proença

Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational complexity and resource requirements. Smaller…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Mahrukh Awan , Asmar Nadeem , Muhammad Junaid Awan , Armin Mustafa , Syed Sameed Husain

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Ya Jiang , Qing Wang , Jun Du , Maocheng Hu , Pengfei Hu , Zeyan Liu , Shi Cheng , Zhaoxu Nian , Yuxuan Dong , Mingqi Cai , Xin Fang , Chin-Hui Lee

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Okan Köpüklü , Maja Taseska , Gerhard Rigoll