中文
相关论文

相关论文: Low-Latency Online Speaker Diarization with Graph-…

200 篇论文

Due to the noises in crowdsourced labels, label aggregation (LA) has emerged as a standard procedure to post-process crowdsourced labels. LA methods estimate true labels from crowdsourced labels by modeling worker qualities. Most existing…

人机交互 · 计算机科学 2022-12-02 Yi Yang , Zhong-Qiu Zhao , Quan Bai , Qing Liu , Weihua Li

Significant progress has recently been made in speaker diarisation after the introduction of d-vectors as speaker embeddings extracted from neural network (NN) speaker classifiers for clustering speech segments. To extract better-performing…

声音 · 计算机科学 2021-05-10 Guangzhi Sun , Chao Zhang , Phil Woodland

This paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering…

With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while deep learning has…

声音 · 计算机科学 2025-09-05 Zhaorui Sun , Yihao Chen , Jialong Wang , Minqiang Xu , Lei Fang , Sian Fang , Lin Liu

Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural network-based generative methods for speaker diarization…

声音 · 计算机科学 2024-09-20 Zhengyang Chen , Bing Han , Shuai Wang , Yidi Jiang , Yanmin Qian

Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming applications. To…

音频与语音处理 · 电气工程与系统科学 2024-11-04 Waris Quamer , Ricardo Gutierrez-Osuna

We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations. Our method formally connects density-based…

机器学习 · 计算机科学 2025-08-06 Ninh Pham , Yingtao Zheng , Hugo Phibbs

This paper describes an effective unsupervised speaker indexing approach. We suggest a two stage algorithm to speed-up the state-of-the-art algorithm based on the Bayesian Information Criterion (BIC). In the first stage of the merging…

声音 · 计算机科学 2010-09-27 Konstantin Biatov

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech…

音频与语音处理 · 电气工程与系统科学 2022-02-07 Naijun Zheng , Na Li , Xixin Wu , Lingwei Meng , Jiawen Kang , Haibin Wu , Chao Weng , Dan Su , Helen Meng

Speaker diarization is the process of labeling different speakers in a speech signal. Deep speaker embeddings are generally extracted from short speech segments and clustered to determine the segments belong to same speaker identity. The…

音频与语音处理 · 电气工程与系统科学 2021-05-18 Myungjong Kim , Vijendra Raj Apsingekar , Divya Neelagiri

We propose a new speaker diarization system based on a recently introduced unsupervised clustering technique namely, generative adversarial network mixture model (GANMM). The proposed system uses x-vectors as front-end representation.…

音频与语音处理 · 电气工程与系统科学 2019-10-28 Monisankha Pal , Manoj Kumar , Raghuveer Peri , Shrikanth Narayanan

Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are…

音频与语音处理 · 电气工程与系统科学 2023-04-19 Francisco Teixeira , Alberto Abad , Bhiksha Raj , Isabel Trancoso

Speaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice profiles. In this paper, we propose to solve the speaker…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Jixuan Wang , Xiong Xiao , Jian Wu , Ranjani Ramamurthy , Frank Rudzicz , Michael Brudno

Speaker diarization answers the question "who spoke when" for an audio file. In some diarization scenarios, low latency is required for transcription. Speaker diarization with low latency is referred to as online speaker diarization. The…

声音 · 计算机科学 2024-08-06 Roman Aperdannier , Sigurd Schacht , Alexander Piazza

Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize large-scale…

音频与语音处理 · 电气工程与系统科学 2023-07-18 Shilong Wu , Jun Du , Maokui He , Shutong Niu , Hang Chen , Haitao Tang , Chin-Hui Lee

We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking…

音频与语音处理 · 电气工程与系统科学 2024-02-09 Srinath Tankasala , Long Chen , Andreas Stolcke , Anirudh Raju , Qianli Deng , Chander Chandak , Aparna Khare , Roland Maas , Venkatesh Ravichandran

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

声音 · 计算机科学 2026-05-11 Yassin Terraf , Youssef Iraqi

This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation method with a traditional multi-module clustering technique to…

声音 · 计算机科学 2025-05-29 Shangkun Huang , Yuxuan Du , Jingwen Yang , Dejun Zhang , Xupeng Jia , Jing Deng , Jintao Kang , Rong Zheng

The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in architecture, objectives, and mixture simulation methods, it…

音频与语音处理 · 电气工程与系统科学 2024-01-30 Desh Raj , Matthew Wiesner , Matthew Maciejewski , Leibny Paola Garcia-Perera , Daniel Povey , Sanjeev Khudanpur

The paper presents a novel approach to refining similarity scores between input utterances for robust speaker verification. Given the embeddings from a pair of input utterances, a graph model is designed to incorporate additional…

音频与语音处理 · 电气工程与系统科学 2021-09-21 Jingyu Li , Si-Ioi Ng , Tan Lee