English
Related papers

Related papers: Leveraging Visual Supervision for Array-based Acti…

200 papers

In this paper, we study teacher-student learning from the perspective of data initialization and propose a novel algorithm called Active Teacher(Source code are available at: \url{https://github.com/HunterJ-Lin/ActiveTeacher}) for…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Peng Mi , Jianghang Lin , Yiyi Zhou , Yunhang Shen , Gen Luo , Xiaoshuai Sun , Liujuan Cao , Rongrong Fu , Qiang Xu , Rongrong Ji

In self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn't always benefit from pseudo labels due to their unreliability. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-15 Ruijie Tao , Kong Aik Lee , Rohan Kumar Das , Ville Hautamäki , Haizhou Li

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based models have shown good performance in this task. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-07 Holger Severin Bovbjerg , Jan Østergaard , Jesper Jensen , Zheng-Hua Tan

State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Théo Lepage , Réda Dehak

Voice activity detection (VAD) is an essential pre-processing step for tasks such as automatic speech recognition (ASR) and speaker recognition. A basic goal is to remove silent segments within an audio, while a more general VAD system…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-22 Yefei Chen , Shuai Wang , Yanmin Qian , Kai Yu

We introduce a distinctive real-time, causal, neural network-based active speaker detection system optimized for low-power edge computing. This system drives a virtual cinematography module and is deployed on a commercial device. The system…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Ilya Gurvich , Ido Leichter , Dharmendar Reddy Palle , Yossi Asher , Alon Vinnikov , Igor Abramovski , Vishak Gopal , Ross Cutler , Eyal Krupka

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

The use of deep networks to extract embeddings for speaker recognition has proven successfully. However, such embeddings are susceptible to performance degradation due to the mismatches among the training, enrollment, and test conditions.…

Sound · Computer Science 2019-04-30 Zhong Meng , Yong Zhao , Jinyu Li , Yifan Gong

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR…

Computer Vision and Pattern Recognition · Computer Science 2017-11-01 Amirsina Torfi , Seyed Mehdi Iranmanesh , Nasser M. Nasrabadi , Jeremy Dawson

Speaker representation learning is crucial for voice recognition systems, with recent advances in self-supervised approaches reducing dependency on labeled data. Current two-stage iterative frameworks, while effective, suffer from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Danwei Cai , Zexin Cai , Ze Li , Ming Li

We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately and do not…

Computer Vision and Pattern Recognition · Computer Science 2021-08-06 Yuanhang Zhang , Susan Liang , Shuang Yang , Xiao Liu , Zhongqin Wu , Shiguang Shan , Xilin Chen

Audio-visual speech contains synchronized audio and visual information that provides cross-modal supervision to learn representations for both automatic speech recognition (ASR) and visual speech recognition (VSR). We introduce continuous…

Machine Learning · Computer Science 2023-10-02 Andrew Rouditchenko , Ronan Collobert , Tatiana Likhomanenko

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persisting challenge. In this paper, we propose a new self-supervised speaker verification approach, Self-Distillation…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-28 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Shiliang Zhang , Wen Wang

Speaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker. This paper demonstrates the effectiveness of meta-learning and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-25 Ashutosh Chaubey , Sparsh Sinha , Susmita Ghose

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also benefit speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-01 Ming Cheng , Ming Li

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Xuan Shi , Erica Cooper , Junichi Yamagishi

Automatic speech recognition (ASR) training can utilize multiple experts as teacher models, each trained on a specific domain or accent. Teacher models may be opaque in nature since their architecture may be not be known or their training…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-29 Aakriti Agrawal , Milind Rao , Anit Kumar Sahu , Gopinath Chennupati , Andreas Stolcke

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listener's brainwaves to a representation of all the sound sources…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-14 Cong Han , Vishal Choudhari , Yinghao Aaron Li , Nima Mesgarani

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to…

Sound · Computer Science 2019-07-03 Miquel India , Pooyan Safari , Javier Hernando