English
Related papers

Related papers: Pushing the Limits of End-to-End Diarization

200 papers

We propose three regularization-based speaker adaptation approaches to adapt the attention-based encoder-decoder (AED) model with very limited adaptation data from target speakers for end-to-end automatic speech recognition. The first…

Computation and Language · Computer Science 2019-11-12 Zhong Meng , Yashesh Gaur , Jinyu Li , Yifan Gong

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fixed (and known) number of speakers, which limits its…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-10 Maokui He , Desh Raj , Zili Huang , Jun Du , Zhuo Chen , Shinji Watanabe

DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker diarization was…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Neville Ryant , Prachi Singh , Venkat Krishnamohan , Rajat Varma , Kenneth Church , Christopher Cieri , Jun Du , Sriram Ganapathy , Mark Liberman

Deep clustering is a recently introduced deep learning architecture that uses discriminatively trained embeddings as the basis for clustering. It was recently applied to spectrogram segmentation, resulting in impressive results on…

Machine Learning · Computer Science 2016-07-11 Yusuf Isik , Jonathan Le Roux , Zhuo Chen , Shinji Watanabe , John R. Hershey

In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, and overlapped speech detection (OSD) plus handling still…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Petr Pálka , Federico Landini , Dominik Klement , Mireia Diez , Anna Silnova , Marc Delcroix , Lukáš Burget

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data…

Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise,…

Sound · Computer Science 2024-06-26 Anurag Chowdhury , Abhinav Misra , Mark C. Fuhs , Monika Woszczyna

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

This paper describes a method for overlap-aware speaker diarization. Given an overlap detector and a speaker embedding extractor, our method performs spectral clustering of segments informed by the output of the overlap detector. This is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Desh Raj , Zili Huang , Sanjeev Khudanpur

In diarization, the PLDA is typically used to model an inference structure which assumes the variation in speech segments be induced by various speakers. The speaker variation is then learned from the training data. However, human…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-01 Jiamin Xie , Suzanna Sia , Paola Garcia , Daniel Povey , Sanjeev Khudanpur

State-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on…

Machine Learning · Statistics 2018-11-02 Vivek Sivaraman Narayanaswamy , Jayaraman J. Thiagarajan , Huan Song , Andreas Spanias

This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automatically extract and…

Computation and Language · Computer Science 2019-09-19 Naoyuki Kanda , Shota Horiguchi , Yusuke Fujita , Yawen Xue , Kenji Nagamatsu , Shinji Watanabe

Speaker diarisation systems nowadays use embeddings generated from speech segments in a bottleneck layer, which are needed to be discriminative for unseen speakers. It is well-known that large-margin training can improve the generalisation…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-07 Yassir Fathullah , Chao Zhang , Philip C. Woodland

End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-22 Jiangyu Han , Federico Landini , Johan Rohdin , Anna Silnova , Mireia Diez , Lukas Burget

Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Huanru Henry Mao , Shuyang Li , Julian McAuley , Garrison Cottrell

Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-11 Xianrui Zheng , Chao Zhang , Philip C. Woodland

This paper describes our solution for the Diarization of Speaker and Language in Conversational Environments Challenge (Displace 2023). We used a combination of VAD for finding segfments with speech, Resnet architecture based CNN for…

Computation and Language · Computer Science 2024-06-25 Ali Aliyev

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-04 Samuele Cornell , Jee-weon Jung , Shinji Watanabe , Stefano Squartini

Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-10 Youngki Kwon , Hee-Soo Heo , Bong-Jin Lee , You Jin Kim , Jee-weon Jung

Regularization is important for end-to-end speech models, since the models are highly flexible and easy to overfit. Data augmentation and dropout has been important for improving end-to-end models in other domains. However, they are…

Computation and Language · Computer Science 2017-12-20 Yingbo Zhou , Caiming Xiong , Richard Socher