中文
相关论文

相关论文: Low-Latency Online Speaker Diarization with Graph-…

200 篇论文

Speaker identification (SID) in the household scenario (e.g., for smart speakers) is an important but challenging problem due to limited number of labeled (enrollment) utterances, confusable voices, and demographic imbalances. Conventional…

音频与语音处理 · 电气工程与系统科学 2023-07-19 Long Chen , Yixiong Meng , Venkatesh Ravichandran , Andreas Stolcke

Recently, there has been increasing progress in end-to-end automatic speech recognition (ASR) architecture, which transcribes speech to text without any pre-trained alignments. One popular end-to-end approach is the hybrid Connectionist…

音频与语音处理 · 电气工程与系统科学 2023-07-06 Haoran Miao , Gaofeng Cheng , Pengyuan Zhang , Yonghong Yan

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (VMFMM) for…

音频与语音处理 · 电气工程与系统科学 2025-02-25 Tobias Cord-Landwehr , Christoph Boeddeker , Reinhold Haeb-Umbach

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

The recently proposed VBx diarization method uses a Bayesian hidden Markov model to find speaker clusters in a sequence of x-vectors. In this work we perform an extensive comparison of performance of the VBx diarization with other…

音频与语音处理 · 电气工程与系统科学 2021-01-01 Federico Landini , Ján Profant , Mireia Diez , Lukáš Burget

We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based EEND model with WavLM features to infer frame-level speaker…

音频与语音处理 · 电气工程与系统科学 2025-10-23 Petr Pálka , Jiangyu Han , Marc Delcroix , Naohiro Tawara , Lukáš Burget

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. This paper proposes a hierarchical network with transformer encoders and memory mechanism to address this problem. The proposed…

声音 · 计算机科学 2020-11-02 Yanpei Shi , Mingjie Chen , Qiang Huang , Thomas Hain

Peer-Led Team Learning (PLTL) is a structured learning model where a team leader is appointed to facilitate collaborative problem solving among students for Science, Technology, Engineering and Mathematics (STEM) courses. This paper…

声音 · 计算机科学 2016-09-28 Harishchandra Dubey , Abhijeet Sangwan , John H. L. Hansen

Getting a robust time-series clustering with best choice of distance measure and appropriate representation is always a challenge. We propose a novel mechanism to identify the clusters combining learned compact representation of…

机器学习 · 计算机科学 2021-01-12 Soma Bandyopadhyay , Anish Datta , Arpan Pal

This paper describes the TSUP team's submission to the ISCSLP 2022 conversational short-phrase speaker diarization (CSSD) challenge which particularly focuses on short-phrase conversations with a new evaluation metric called conversational…

声音 · 计算机科学 2023-10-26 Bowen Pang , Huan Zhao , Gaosheng Zhang , Xiaoyue Yang , Yang Sun , Li Zhang , Qing Wang , Lei Xie

We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a configuration file. Such flexibility allows our system to work…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Giovanni Morrone , Enrico Zovato , Fabio Brugnara , Enrico Sartori , Leonardo Badino

Speaker diarization determines who spoke and when? in an audio stream. In this study, we propose a model-based approach for robust speaker clustering using i-vectors. The ivectors extracted from different segments of same speaker are…

声音 · 计算机科学 2019-07-15 Harishchandra Dubey , Abhijeet Sangwan , John Hansen

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be…

音频与语音处理 · 电气工程与系统科学 2020-05-26 Jixuan Wang , Xiong Xiao , Jian Wu , Ranjani Ramamurthy , Frank Rudzicz , Michael Brudno

Speaker diarization based on bottom-up clustering of speech segments by acoustic similarity is often highly sensitive to the choice of hyperparameters, such as the initial number of clusters and feature weighting. Optimizing these…

计算与语言 · 计算机科学 2022-02-22 Andreas Stolcke

This paper describes the systems submitted by team HCCL to the Far-Field Speaker Verification Challenge. Our previous work in the AIshell Speaker Verification Challenge 2019 shows that the powerful modeling abilities of Neural Network…

声音 · 计算机科学 2021-07-06 Zhuo Li , Ce Fang , Runqiu Xiao , Zhigao Chen , Wenchao Wang , Yonghong Yan

An enhanced label propagation (LP) method called GraphHop was proposed recently. It outperforms graph convolutional networks (GCNs) in the semi-supervised node classification task on various networks. Although the performance of GraphHop…

机器学习 · 计算机科学 2022-11-01 Tian Xie , Rajgopal Kannan , C. -C. Jay Kuo

In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial…

音频与语音处理 · 电气工程与系统科学 2019-11-05 Raghuveer Peri , Monisankha Pal , Arindam Jati , Krishna Somandepalli , Shrikanth Narayanan

In this study, we propose a new spectral clustering framework that can auto-tune the parameters of the clustering algorithm in the context of speaker diarization. The proposed framework uses normalized maximum eigengap (NME) values to…

音频与语音处理 · 电气工程与系统科学 2020-06-29 Tae Jin Park , Kyu J. Han , Manoj Kumar , Shrikanth Narayanan

Speaker clustering is the task of forming speaker-specific groups based on a set of utterances. In this paper, we address this task by using Dominant Sets (DS). DS is a graph-based clustering algorithm with interesting properties that fits…

声音 · 计算机科学 2018-05-23 Feliks Hibraj , Sebastiano Vascon , Thilo Stadelmann , Marcello Pelillo

In this paper, we propose an online speaker diarization system based on Relation Network, named RenoSD. Unlike conventional diariztion systems which consist of several independently-optimized modules, RenoSD implements…

音频与语音处理 · 电气工程与系统科学 2020-09-22 Xiang Li , Yucheng Zhao , Chong Luo , Wenjun Zeng