中文
相关论文

相关论文: Learnable MFCCs for Speaker Verification

200 篇论文

Speech enhancement has benefited from the success of deep learning in terms of intelligibility and perceptual quality. Conventional time-frequency (TF) domain methods focus on predicting TF-masks or speech spectrum, via a naive convolution…

音频与语音处理 · 电气工程与系统科学 2020-09-24 Yanxin Hu , Yun Liu , Shubo Lv , Mengtao Xing , Shimin Zhang , Yihui Fu , Jian Wu , Bihong Zhang , Lei Xie

There are many deterministic mathematical operations (e.g. compression, clipping, downsampling) that degrade speech quality considerably. In this paper we introduce a neural network architecture, based on a modification of the DiffWave…

声音 · 计算机科学 2021-09-03 Jianwei Zhang , Suren Jayasuriya , Visar Berisha

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic…

声音 · 计算机科学 2018-06-15 Yi Liu , Liang He , Jia Liu , Michael T. Johnson

Most speaker verification tasks are studied as an open-set evaluation scenario considering the real-world condition. Thus, the generalization power to unseen speakers is of paramount important to the performance of the speaker verification…

音频与语音处理 · 电气工程与系统科学 2021-04-15 Ju-ho Kim , Hye-jin Shim , Jee-weon Jung , Ha-Jin Yu

Deep learning has enabled highly realistic synthetic speech, raising concerns about fraud, impersonation, and disinformation. Despite rapid progress in neural detectors, transparent baselines are needed to reveal which acoustic cues…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Faheem Ahmad , Ajan Ahmed , Masudul Imtiaz

Cardiovascular system diseases can be identified by using a specialized diagnostic process utilizing a digital stethoscope. Digital stethoscopes provide phonocardiography (PCG) recordings for further inspection, besides filtering and…

信号处理 · 电气工程与系统科学 2024-02-21 Ibrahim Ozkan , Atila Yilmaz

Time delay neural network (TDNN) has been proven to be efficient for speaker verification. One of its successful variants, ECAPA-TDNN, achieved state-of-the-art performance at the cost of much higher computational complexity and slower…

声音 · 计算机科学 2023-06-19 Hui Wang , Siqi Zheng , Yafeng Chen , Luyao Cheng , Qian Chen

Fast Fourier convolution (FFC) is the recently proposed neural operator showing promising performance in several computer vision problems. The FFC operator allows employing large receptive field operations within early layers of the neural…

声音 · 计算机科学 2022-04-08 Ivan Shchekotov , Pavel Andreev , Oleg Ivanov , Aibek Alanov , Dmitry Vetrov

This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and encoding of the input text. We showed that the latent space…

音频与语音处理 · 电气工程与系统科学 2019-07-11 Daniel Korzekwa , Roberto Barra-Chicote , Bozena Kostek , Thomas Drugman , Mateusz Lajszczak

Neural network based speech dereverberation has achieved promising results in recent studies. Nevertheless, many are focused on recovery of only the direct path sound and early reflections, which could be beneficial to speech perception,…

声音 · 计算机科学 2021-10-19 Ziteng Wang , Yueyue Na , Biao Tian , Qiang Fu

Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based…

声音 · 计算机科学 2024-09-04 Tathagata Bandyopadhyay

Over the recent years, various deep learning-based embedding methods have been proposed and have shown impressive performance in speaker verification. However, as in most of the classical embedding techniques, the deep learning-based…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Woo Hyun Kang , Sung Hwan Mun , Min Hyun Han , Nam Soo Kim

The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical…

In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly…

声音 · 计算机科学 2025-05-07 Ya Li , Bin Zhou , Bo Hu

LSTM-based speaker verification usually uses a fixed-length local segment randomly truncated from an utterance to learn the utterance-level speaker embedding, while using the average embedding of all segments of a test utterance to verify…

音频与语音处理 · 电气工程与系统科学 2018-11-05 Bin Liu , Shuai Nie , Yaping Zhang , Shan Liang , Wenju Liu

Traditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values. Recent studies found that AMs with convolutional neural networks…

音频与语音处理 · 电气工程与系统科学 2019-10-10 Patrick von Platen , Chao Zhang , Philip Woodland

Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Jingyu Li , Yusheng Tian , Tan Lee

Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its…

声音 · 计算机科学 2024-07-11 Guoqiang Hu , Huaning Tan , Ruilai Li

In this work, we introduce metric learning (ML) to enhance the deep embedding learning for text-independent speaker verification (SV). Specifically, the deep speaker embedding network is trained with conventional cross entropy loss and…

音频与语音处理 · 电气工程与系统科学 2023-03-24 Yafeng Chen , Wu Guo , Jingjing Shi , Jiajun Qi , Tan Liu

Deep Convolutional Neural Networks (CNNs) are more powerful than Deep Neural Networks (DNN), as they are able to better reduce spectral variation in the input signal. This has also been confirmed experimentally, with CNNs showing…

‹ 上一页 1 8 9 10 下一页 ›