中文
相关论文

相关论文: SpeakerNet: 1D Depth-wise Separable Convolutional …

200 篇论文

In this paper, we propose an iterative framework for self-supervised speaker representation learning based on a deep neural network (DNN). The framework starts with training a self-supervision speaker embedding network by maximizing…

音频与语音处理 · 电气工程与系统科学 2020-10-29 Danwei Cai , Weiqing Wang , Ming Li

State-of-the-art speaker verification models are based on deep learning techniques, which heavily depend on the handdesigned neural architectures from experts or engineers. We borrow the idea of neural architecture search(NAS) for the…

音频与语音处理 · 电气工程与系统科学 2020-08-14 Xiaoyang Qu , Jianzong Wang , Jing Xiao

This work considers training neural networks for speaker recognition with a much smaller dataset size compared to contemporary work. We artificially restrict the amount of data by proposing three subsets of the popular VoxCeleb2 dataset.…

声音 · 计算机科学 2023-02-28 Nik Vaessen , David A. van Leeuwen

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is provided only at the…

声音 · 计算机科学 2018-06-25 Martin Karu , Tanel Alumäe

This work presents a framework based on feature disentanglement to learn speaker embeddings that are robust to environmental variations. Our framework utilises an auto-encoder as a disentangler, dividing the input speaker embedding into…

声音 · 计算机科学 2024-06-21 KiHyun Nam , Hee-Soo Heo , Jee-weon Jung , Joon Son Chung

This paper summarizes the applied deep learning practices in the field of speaker recognition, both verification and identification. Speaker recognition has been a widely used field topic of speech technology. Many research works have been…

音频与语音处理 · 电气工程与系统科学 2022-09-27 Dávid Sztahó , György Szaszák , András Beke

Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative…

声音 · 计算机科学 2021-10-14 Xiaoyi Qin , Na Li , Chao Weng , Dan Su , Ming Li

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker…

音频与语音处理 · 电气工程与系统科学 2023-11-02 Tianchi Liu , Kong Aik Lee , Qiongqiong Wang , Haizhou Li

This paper describes the IDLab submission for the text-independent task of the Short-duration Speaker Verification Challenge 2021 (SdSVC-21). This speaker verification competition focuses on short duration test recordings and cross-lingual…

音频与语音处理 · 电气工程与系统科学 2021-09-10 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residual network (ResNet)…

音频与语音处理 · 电气工程与系统科学 2019-12-30 Youngmoon Jung , Younggwan Kim , Hyungjun Lim , Yeunju Choi , Hoirin Kim

This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequence architecture of our previous target-speaker voice…

音频与语音处理 · 电气工程与系统科学 2025-06-24 Ming Cheng , Yuke Lin , Ming Li

Today, Time Delay Neural Network (TDNN) has become the mainstream architecture for speaker verification task, in which the ECAPA-TDNN is one of the state-of-the-art models. The current works that focus on improving TDNN primarily address…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Shilong Weng , Liu Yang , Ji Mao

This work presents a novel back-end framework for speaker verification using graph attention networks. Segment-wise speaker embeddings extracted from multiple crops within an utterance are interpreted as node representations of a graph. The…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Jee-weon Jung , Hee-Soo Heo , Ha-Jin Yu , Joon Son Chung

Automatic speaker recognition algorithms typically use pre-defined filterbanks, such as Mel-Frequency and Gammatone filterbanks, for characterizing speech audio. However, it has been observed that the features extracted using these…

音频与语音处理 · 电气工程与系统科学 2022-06-14 Anurag Chowdhury , Arun Ross

In this paper, we propose an online speaker adaptation method for WaveNet-based neural vocoders in order to improve their performance on speaker-independent waveform generation. In this method, a speaker encoder is first constructed using a…

音频与语音处理 · 电气工程与系统科学 2020-08-17 Qiuchen Huang , Yang Ai , Zhenhua Ling

Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on implicit noise suppression, which struggles to separate noise…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Minu Kim , Kangwook Jang , Hoirin Kim

The task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene…

音频与语音处理 · 电气工程与系统科学 2019-11-05 Fabian-Robert Stöter , Soumitro Chakrabarty , Bernd Edler , Emanuël A. P. Habets

The classical i-vectors and the latest end-to-end deep speaker embeddings are the two representative categories of utterance-level representations in automatic speaker verification systems. Traditionally, once i-vectors or deep speaker…

音频与语音处理 · 电气工程与系统科学 2018-06-12 Weicheng Cai , Jinkun Chen , Ming Li

Speaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker. This paper demonstrates the effectiveness of meta-learning and…

音频与语音处理 · 电气工程与系统科学 2022-07-25 Ashutosh Chaubey , Sparsh Sinha , Susmita Ghose

Deep learning approaches are still not very common in the speaker verification field. We investigate the possibility of using deep residual convolutional neural network with spectrograms as an input features in the text-dependent speaker…

声音 · 计算机科学 2017-05-31 Egor Malykh , Sergey Novoselov , Oleg Kudashev