English
Related papers

Related papers: End-to-End Neural Diarization: Reformulating Speak…

200 papers

Automatic speaker diarization techniques typically involve a two-stage processing approach where audio segments of fixed duration are converted to vector representations in the first stage. This is followed by an unsupervised clustering of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Prachi Singh , Sriram Ganapathy

We address the problem of acoustic source separation in a deep learning framework we call "deep clustering." Rather than directly estimating signals or masking functions, we train a deep network to produce spectrogram embeddings that are…

Neural and Evolutionary Computing · Computer Science 2015-08-19 John R. Hershey , Zhuo Chen , Jonathan Le Roux , Shinji Watanabe

State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer to encode…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-08 Federico Costa , Miquel India , Javier Hernando

Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering…

Computation and Language · Computer Science 2019-02-08 Yiming Wang , Xing Fan , I-Fan Chen , Yuzong Liu , Tongfei Chen , Björn Hoffmeister

Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings…

Sound · Computer Science 2021-02-15 G. Sun , D. Liu , C. Zhang , P. C. Woodland

Accurate transcription and speaker diarization of child-adult spoken interactions are crucial for developmental and clinical research. However, manual annotation is time-consuming and challenging to scale. Existing automated systems…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Anfeng Xu , Tiantian Feng , Somer Bishop , Catherine Lord , Shrikanth Narayanan

DER is the primary metric to evaluate diarization performance while facing a dilemma: the errors in short utterances or segments tend to be overwhelmed by longer ones. Short segments, e.g., `yes' or `no,' still have semantic information.…

Sound · Computer Science 2022-11-09 Tao Liu , Kai Yu

This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automatically extract and…

Computation and Language · Computer Science 2019-09-19 Naoyuki Kanda , Shota Horiguchi , Yusuke Fujita , Yawen Xue , Kenji Nagamatsu , Shinji Watanabe

Many modern systems for speaker diarization, such as the recently-developed VBx approach, rely on clustering of DNN speaker embeddings followed by resegmentation. Two problems with this approach are that the DNN is not directly optimized…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Kiran Karra , Alan McCree

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

Sound · Computer Science 2021-04-26 Chau Luu , Peter Bell , Steve Renals

Voice activity detection (VAD) is essential in speech-based systems, but traditional methods detect only speech presence without identifying speakers. Target-speaker VAD (TS-VAD) extends this by detecting the speech of a known speaker using…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-16 Wen-Yung Wu , Pei-Chin Hsieh , Tai-Shih Chi

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…

Sound · Computer Science 2022-11-04 You Jin Kim , Hee-Soo Heo , Jee-weon Jung , Youngki Kwon , Bong-Jin Lee , Joon Son Chung

We propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization. Our proposed model converts song data with choral singing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-25 Hokuto Munakata , Ryo Terashima , Yusuke Fujita

Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture, while speaker diarization demarcates speech segments by…

Sound · Computer Science 2025-01-17 Junyi Ao , Mehmet Sinan Yıldırım , Ruijie Tao , Meng Ge , Shuai Wang , Yanmin Qian , Haizhou Li

End-to-end learning treats the entire system as a whole adaptable black box, which, if sufficient data are available, may learn a system that works very well for the target task. This principle has recently been applied to several prototype…

Sound · Computer Science 2017-06-27 Dong Wang , Lantian Li , Zhiyuan Tang , Thomas Fang Zheng

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-14 Weiqing Wang , Qingjian Lin , Ming Li

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as feature extractors for speaker verification has shown…

Computation and Language · Computer Science 2017-01-04 Shi-Xiong Zhang , Zhuo Chen , Yong Zhao , Jinyu Li , Yifan Gong

Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Taous Iatariene , Alexandre Guérin , Romain Serizel

Recently, an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identification for monaural overlapped speech. It showed…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Naoyuki Kanda , Xuankai Chang , Yashesh Gaur , Xiaofei Wang , Zhong Meng , Zhuo Chen , Takuya Yoshioka

We provide the technical report for Ego4D audio-only diarization challenge in ECCV 2022. Speaker diarization takes the audio streams as input and outputs the homogeneous segments according to the speaker's identity. It aims to solve the…

Sound · Computer Science 2022-11-17 Jiahao Wang , Guo Chen , Yin-Dong Zheng , Tong Lu