English
Related papers

Related papers: Target Active Speaker Detection with Audio-visual …

200 papers

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

Computation and Language · Computer Science 2023-05-15 Fei Tao , Carlos Busso

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yidi Li , Hong Liu , Bing Yang

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Yidi Jiang , Ruijie Tao , Zhengyang Chen , Yanmin Qian , Haizhou Li

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. Building upon the achievements of the state-of-the-art (SOTA)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-31 Zexu Pan , Gordon Wichern , Yoshiki Masuyama , Francois G. Germain , Sameer Khurana , Chiori Hori , Jonathan Le Roux

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Target-Speaker Voice Activity Detection (TS-VAD) utilizes a set of speaker profiles alongside an input audio signal to perform speaker diarization. While its superiority over conventional methods has been demonstrated, the method can suffer…

Sound · Computer Science 2024-04-05 Dongmei Wang , Xiong Xiao , Naoyuki Kanda , Midia Yousefi , Takuya Yoshioka , Jian Wu

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Voice activity detection (VAD) is essential in speech-based systems, but traditional methods detect only speech presence without identifying speakers. Target-speaker VAD (TS-VAD) extends this by detecting the speech of a known speaker using…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-16 Wen-Yung Wu , Pei-Chin Hsieh , Tai-Shih Chi

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Ruijie Tao , Rohan Kumar Das , Haizhou Li

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Zexu Pan , Shengkui Zhao , Tingting Wang , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma

Active speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction. This paper introduces FabuLight-ASD, an advanced ASD model that integrates facial, audio, and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Hugo Carneiro , Stefan Wermter

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-10 Zexu Pan , Wupeng Wang , Marvin Borsdorf , Haizhou Li

Humans exhibit a remarkable ability to focus auditory attention in complex acoustic environments, such as cocktail parties. Auditory attention detection (AAD) aims to identify the attended speaker by analyzing brain signals, such as…

Signal Processing · Electrical Eng. & Systems 2025-03-07 Yuan Liao , Yuhong Zhang , Qiushi Han , Yuhang Yang , Weiwei Ding , Yuzhe Gu , Hengxin Yang , Liya Huang

Speaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target speaker extraction, e.g., SpEx, we propose a unified speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-05 Chenglin Xu , Wei Rao , Jibin Wu , Haizhou Li

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting information about a…

The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error loss between the DNN's…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Yuewei Zhang , Huanbin Zou , Jie Zhu

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-14 Weiqing Wang , Qingjian Lin , Ming Li

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

This report describes our submission to the ActivityNet Challenge at CVPR 2019. We use a 3D convolutional neural network (CNN) based front-end and an ensemble of temporal convolution and LSTM classifiers to predict whether a visible person…

Sound · Computer Science 2019-06-26 Joon Son Chung

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu