English
Related papers

Related papers: Target Active Speaker Detection with Audio-visual …

200 papers

Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-12 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Xinyue Song , Xianghu Yue

This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV).…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-17 Chunlei Zhang , Meng Yu , Chao Weng , Dong Yu

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Kevin Huang , Tiantian Feng , Helen Tager-Flusberg , Shrikanth Narayanan

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The classifier uses…

Computer Vision and Pattern Recognition · Computer Science 2016-03-30 Punarjay Chakravarty , Tinne Tuytelaars

Extracting the speech of a target speaker from mixed audios, based on a reference speech from the target speaker, is a challenging yet powerful technology in speech processing. Recent studies of speaker-independent speech separation, such…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Zining Zhang , Bingsheng He , Zhenjie Zhang

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

This study considers the problem of detecting and locating an active talker's horizontal position from multichannel audio captured by a microphone array. We refer to this as active speaker detection and localization (ASDL). Our goal was to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Davide Berghi , Philip J. B. Jackson

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

Sound · Computer Science 2026-05-11 Yassin Terraf , Youssef Iraqi

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

Multimedia · Computer Science 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

This paper presents a streaming speaker-attributed automatic speech recognition (SA-ASR) model that can recognize ``who spoke what'' with low latency even when multiple people are speaking simultaneously. Our model is based on token-level…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-18 Naoyuki Kanda , Jian Wu , Yu Wu , Xiong Xiao , Zhong Meng , Xiaofei Wang , Yashesh Gaur , Zhuo Chen , Jinyu Li , Takuya Yoshioka

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected speaker not…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Tomoya Yoshinaga , Keitaro Tanaka , Shigeo Morishima

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Auditory spatial attention detection (ASAD) is used to determine the direction of a listener's attention to a speaker by analyzing her/his electroencephalographic (EEG) signals. This study aimed to further improve the performance of ASAD…

Signal Processing · Electrical Eng. & Systems 2024-05-15 Yuting Ding , Fei Chen

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Yuanhang Zhang , Susan Liang , Shuang Yang , Shiguang Shan