English
Related papers

Related papers: On Attention Modules for Audio-Visual Synchronizat…

200 papers

Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Fei Yu , Jun Wang

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

We study cross-modal recommendation of music tracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association between music and video.…

Multimedia · Computer Science 2023-06-13 Laure Prétet , Gaël Richard , Clément Souchier , Geoffroy Peeters

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

Sound · Computer Science 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-28 Nasir Saleem , Mandar Gogate , Kia Dashtipour , Adeel Hussain , Usman Anwar , Adewale Adetomi , Tughrul Arslan , Amir Hussain

In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This review explores…

Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jaewoo Lee , Jaehong Yoon , Wonjae Kim , Yunji Kim , Sung Ju Hwang

In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design attentional models that can dynamically focus on parts of images…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Loris Bazzani , Hugo Larochelle , Lorenzo Torresani

Performance-score synchronization is an integral task in signal processing, which entails generating an accurate mapping between an audio recording of a performance and the corresponding musical score. Traditional synchronization methods…

Sound · Computer Science 2022-04-20 Ruchit Agrawal , Daniel Wolff , Simon Dixon

Recently, substantial research effort has focused on how to apply CNNs or RNNs to better extract temporal patterns from videos, so as to improve the accuracy of video classification. In this paper, however, we show that temporal…

Computer Vision and Pattern Recognition · Computer Science 2017-11-28 Xiang Long , Chuang Gan , Gerard de Melo , Jiajun Wu , Xiao Liu , Shilei Wen

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

Inspired by the observation that humans are able to process videos efficiently by only paying attention where and when it is needed, we propose an interpretable and easy plug-in spatial-temporal attention mechanism for video action…

Computer Vision and Pattern Recognition · Computer Science 2019-06-04 Lili Meng , Bo Zhao , Bo Chang , Gao Huang , Wei Sun , Frederich Tung , Leonid Sigal

In this paper, we introduce a new problem, named audio-visual video parsing, which aims to parse a video into temporal event segments and label them as either audible, visible, or both. Such a problem is essential for a complete…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Yapeng Tian , Dingzeyu Li , Chenliang Xu

Immersive audio-visual perception relies on the spatial integration of both auditory and visual information which are heterogeneous sensing modalities with different fields of reception and spatial resolution. This study investigates the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-17 Davide Berghi , Hanne Stenzel , Marco Volino , Adrian Hilton , Philip J. B. Jackson

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Ruohan Gao , Kristen Grauman