English
Related papers

Related papers: Variational Bayesian Inference for Audio-Visual Tr…

200 papers

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jun Yu , Gongpeng Zhao , Yongqi Wang , Zhihong Wei , Yang Zheng , Zerui Zhang , Zhongpeng Cai , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Understanding the relationship between the auditory and visual signals is crucial for many different applications ranging from computer-generated imagery (CGI) and video editing automation to assisting people with hearing or visual…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

This paper considers the problem of multiple human target tracking in a sequence of video data. A solution is proposed which is able to deal with the challenges of a varying number of targets, interactions and when every target gives rise…

Computer Vision and Pattern Recognition · Computer Science 2015-11-06 Ata-ur-Rehman , Syed Mohsen Naqvi , Lyudmila Mihaylova , Jonathon Chambers

Tracking multiple particles in noisy and cluttered scenes remains challenging due to a combinatorial explosion of trajectory hypotheses, which scales super-exponentially with the number of particles and frames. The transformer architecture…

Machine Learning · Statistics 2025-06-12 Piyush Mishra , Philippe Roudot

Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Adrian Meise , Tobias Cord-Landwehr , Christoph Boeddeker , Marc Delcroix , Tomohiro Nakatani , Reinhold Haeb-Umbach

In this work, we propose a novel variational Bayesian adaptive learning approach for cross-domain knowledge transfer to address acoustic mismatches between training and testing conditions, such as recording devices and environmental noise.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-28 Hu Hu , Sabato Marco Siniscalchi , Chao-Han Huck Yang , Chin-Hui Lee

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (VMFMM) for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Tobias Cord-Landwehr , Christoph Boeddeker , Reinhold Haeb-Umbach

We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Xixi Hu , Ziyang Chen , Andrew Owens

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Thanh-Dat Truong , Chi Nhan Duong , The De Vu , Hoang Anh Pham , Bhiksha Raj , Ngan Le , Khoa Luu

Passive monitoring of acoustic or radio sources has important applications in modern convenience, public safety, and surveillance. A key task in passive monitoring is multiobject tracking (MOT). This paper presents a Bayesian method for…

Signal Processing · Electrical Eng. & Systems 2024-02-29 Wenyu Zhang , Florian Meyer

The problem of multimodal clustering arises whenever the data are gathered with several physically different sensors. Observations from different modalities are not necessarily aligned in the sense there there is no obvious way to associate…

Machine Learning · Statistics 2020-12-10 Vasil Khalidov , Florence Forbes , Radu Horaud

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spatially small and may occur only infrequently during a many…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Gwanghyun Kim , Alonso Martinez , Yu-Chuan Su , Brendan Jou , José Lezama , Agrim Gupta , Lijun Yu , Lu Jiang , Aren Jansen , Jacob Walker , Krishna Somandepalli

We propose a diarization system, that estimates "who spoke when" based on spatial information, to be used as a front-end of a meeting transcription system running on the signals gathered from an acoustic sensor network (ASN). Although the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-28 Tobias Gburrek , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate detection when the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Juan Leon Alcazar , Fabian Caba Heilbron , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Ruohan Gao , Kristen Grauman

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face relationship are i)…

Multimedia · Computer Science 2022-12-02 Rahul Sharma , Shrikanth Narayanan
‹ Prev 1 3 4 5 6 7 10 Next ›