English
Related papers

Related papers: Rethinking the visual cues in audio-visual speaker…

200 papers

The objective of this paper is to perform audio-visual sound source separation, i.e.~to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in the input video…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Lingyu Zhu , Esa Rahtu

Visual Speech Recognition (VSR) is the task of predicting spoken words from silent lip movements. VSR is regarded as a challenging task because of the insufficient information on lip movements. In this paper, we propose an Audio Knowledge…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Jeong Hun Yeo , Minsu Kim , Jeongsoo Choi , Dae Hoe Kim , Yong Man Ro

Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent…

Sound · Computer Science 2022-04-19 Jen-Cheng Hou , Syu-Siang Wang , Ying-Hui Lai , Yu Tsao , Hsiu-Wen Chang , Hsin-Min Wang

Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent…

Sound · Computer Science 2018-01-25 Jen-Cheng Hou , Syu-Siang Wang , Ying-Hui Lai , Yu Tsao , Hsiu-Wen Chang , Hsin-Min Wang

As a practical alternative of speech separation, target speaker extraction (TSE) aims to extract the speech from the desired speaker using additional speaker cue extracted from the speaker. Its main challenge lies in how to properly extract…

Sound · Computer Science 2023-01-18 Kai Liu , Xucheng Wan , Ziqing Du , Huan Zhou

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum…

Sound · Computer Science 2026-03-06 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Yihan Wu , Yichen Lu , Yifan Peng , Xihua Wang , Ruihua Song , Shinji Watanabe

Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Jiachen Lian , Alexei Baevski , Wei-Ning Hsu , Michael Auli

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Developing a single-microphone speech denoising or dereverberation front-end for robust automatic speaker verification (ASV) in noisy far-field speaking scenarios is challenging. To address this problem, we present a novel front-end design…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Joon-Young Yang , Joon-Hyuk Chang

This manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since the blind IVE cannot…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-29 Jiri Malek , Jakub Jansky , Zbynek Koldovsky , Tomas Kounovsky , Jaroslav Cmejla , Jindrich Zdansky

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Karren Yang , Dejan Markovic , Steven Krenn , Vasu Agrawal , Alexander Richard

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

Image and Video Processing · Electrical Eng. & Systems 2022-09-27 Rahul Sharma , Shrikanth Narayanan

Recent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of…

Sound · Computer Science 2023-08-16 Jeongsoo Choi , Joanna Hong , Yong Man Ro

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE…

Sound · Computer Science 2025-03-13 Minsu Kim , Rodrigo Mira , Honglie Chen , Stavros Petridis , Maja Pantic

Recent techniques for speech deepfake detection often rely on pre-trained self-supervised models. These systems, initially developed for Automatic Speech Recognition (ASR), have proved their ability to offer a meaningful representation of…

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

Computation and Language · Computer Science 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain