English
Related papers

Related papers: Data Fusion for Audiovisual Speaker Localization: …

200 papers

For recognizing speakers in video streams, significant research studies have been made to obtain a rich machine learning model by extracting high-level speaker's features such as facial expression, emotion, and gender. However, generating…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Ehsan Asali , Farzan Shenavarmasouleh , Farid Ghareh Mohammadi , Prasanth Sengadu Suresh , Hamid R. Arabnia

Acoustic source localization has been applied in different fields, such as aeronautics and ocean science, generally using multiple microphones array data to reconstruct the source location. However, the model-based beamforming methods fail…

Sound · Computer Science 2022-04-01 Guanxing Zhou , Hao Liang , Xinghao Ding , Yue Huang , Xiaotong Tu , Saqlain Abbas

Fusing outputs from automatic speaker verification (ASV) and spoofing countermeasure (CM) is expected to make an integrated system robust to zero-effort imposters and synthesized spoofing attacks. Many score-level fusion methods have been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Xin Wang , Tomi Kinnunen , Kong Aik Lee , Paul-Gauthier Noé , Junichi Yamagishi

Recent years have witnessed the increasing application of place recognition in various environments, such as city roads, large buildings, and a mix of indoor and outdoor places. This task, however, still remains challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-24 Haowen Lai , Peng Yin , Sebastian Scherer

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Combining multiple complementary techniques together has long been regarded as a way to improve performance. In visual localization, multi-sensor fusion, multi-process fusion of a single sensing modality, and even combinations of different…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Stephen Hausler , Michael Milford

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Self-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks. Fine-tuning separate foundation models for many downstream…

Machine Learning · Computer Science 2022-11-08 Zhouyuan Huo , Khe Chai Sim , Bo Li , Dongseong Hwang , Tara N. Sainath , Trevor Strohman

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-09 Yiwen Shao

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

Sound · Computer Science 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zijun Ding , Mingdie Xiong , Congcong Zhu , Jingrun Chen

A common problem in acoustic design is the placement of speakers or receivers for public address systems, telecommunications, and home smart speakers or digital personal assistants. We present a novel algorithm to automatically place a…

Sound · Computer Science 2020-02-11 Nicolas Morales , Zhenyu Tang , Dinesh Manocha

This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for various multi-speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-02 Shota Horiguchi , Atsushi Ando , Takafumi Moriya , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

Several advances have been made recently towards handling overlapping speech for speaker diarization. Since speech and natural language tasks often benefit from ensemble techniques, we propose an algorithm for combining outputs from such…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Desh Raj , Leibny Paola Garcia-Perera , Zili Huang , Shinji Watanabe , Daniel Povey , Andreas Stolcke , Sanjeev Khudanpur

A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks. While there have been several augmentation methods shown to…

Sound · Computer Science 2021-08-09 Gwantae Kim , David K. Han , Hanseok Ko

In this paper, we present a deep neural network-based online multi-speaker localisation algorithm. Following the W-disjoint orthogonality principle in the spectral domain, each time-frequency (TF) bin is dominated by a single speaker, and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Hodaya Hammer , Shlomo E. Chazan , Jacob Goldberger , Sharon Gannot

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Kyle Min , Sourya Roy , Subarna Tripathi , Tanaya Guha , Somdeb Majumdar

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Jaesung Huh , Andrew Zisserman

The media localization industry usually requires a verbatim script of the final film or TV production in order to create subtitles or dubbing scripts in a foreign language. In particular, the verbatim script (i.e. as-broadcast script) must…

Computation and Language · Computer Science 2023-08-07 Yogesh Virkar , Brian Thompson , Rohit Paturi , Sundararajan Srinivasan , Marcello Federico
‹ Prev 1 8 9 10 Next ›