中文
相关论文

相关论文: ReDimNet2: Scaling Speaker Verification via Time-P…

200 篇论文

This paper proposes Remixed2Remixed, a domain adaptation method for speech enhancement, which adopts Noise2Noise (N2N) learning to adapt models trained on artificially generated (out-of-domain: OOD) noisy-clean pair data to better separate…

声音 · 计算机科学 2023-12-29 Li Li , Shogo Seki

Semantic segmentation plays a key role in applications such as autonomous driving and medical image. Although existing real-time semantic segmentation models achieve a commendable balance between accuracy and speed, their multi-path blocks…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Guoyu Yang , Yuan Wang , Daming Shi

In contrast to other sequence tasks modeling hidden layer features with three axes, Dual-Path time and time-frequency domain speech enhancement models are effective and have low parameters but are computationally demanding due to their…

声音 · 计算机科学 2025-01-10 Haoxu Wang , Biao Tian

In this paper, we present a Mirroring Neural Network architecture to perform non-linear dimensionality reduction and Object Recognition using a reduced lowdimensional characteristic vector. In addition to dimensionality reduction, the…

计算机视觉与模式识别 · 计算机科学 2008-12-13 Dasika Ratna Deepthi , Sujeet Kuchibhotla , K. Eswaran

Target speaker extraction (TSE) is essential in speech processing applications, particularly in scenarios with complex acoustic environments. Current TSE systems face challenges in limited data diversity and a lack of robustness in…

声音 · 计算机科学 2024-12-18 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Recent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of…

声音 · 计算机科学 2023-08-16 Jeongsoo Choi , Joanna Hong , Yong Man Ro

Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Umut Isik , Ritwik Giri , Neerad Phansalkar , Jean-Marc Valin , Karim Helwani , Arvindh Krishnaswamy

In this paper, we propose an effective training strategy to ex-tract robust speaker representations from a speech signal. Oneof the key challenges in speaker recognition tasks is to learnlatent representations or embeddings containing…

音频与语音处理 · 电气工程与系统科学 2020-08-05 Yoohwan Kwon , Soo-Whan Chung , Hong-Goo Kang

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Chong Deng , Shiliang Zhang , Wen Wang

We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker classification and using…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Amber Afshan , Abeer Alwan

DNN-based speaker verification (SV) models demonstrate significant performance at relatively high computation costs. Model compression can be applied to reduce the model size for lower resource consumption. The present study exploits weight…

音频与语音处理 · 电气工程与系统科学 2023-09-26 Jingyu Li , Wei Liu , Zhaoyang Zhang , Jiong Wang , Tan Lee

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persisting challenge. In this paper, we propose a new self-supervised speaker verification approach, Self-Distillation…

音频与语音处理 · 电气工程与系统科学 2024-12-28 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Shiliang Zhang , Wen Wang

In this paper, we propose to extend the deep, complex U-Network architecture for speech enhancement by incorporating a probabilistic (i.e., variational) latent space model. The proposed model is evaluated against several ablated versions of…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Eike J. Nustede , Jörn Anemüller

Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbone networks for speaker embedding are designed to aggregate…

声音 · 计算机科学 2022-03-18 Ruiteng Zhang , Jianguo Wei , Xugang Lu , Wenhuan Lu , Di Jin , Junhai Xu , Lin Zhang , Yantao Ji , Jianwu Dang

We present MarbleNet, an end-to-end neural network for Voice Activity Detection (VAD). MarbleNet is a deep residual network composed from blocks of 1D time-channel separable convolution, batch-normalization, ReLU and dropout layers. When…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Fei Jia , Somshubra Majumdar , Boris Ginsburg

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand…

音频与语音处理 · 电气工程与系统科学 2018-09-13 Suwon Shon , Hao Tang , James Glass

This paper describes the NPU system submitted to Spoofing Aware Speaker Verification Challenge 2022. We particularly focus on the \textit{backend ensemble} for speaker verification and spoofing countermeasure from three aspects. Firstly,…

声音 · 计算机科学 2022-09-26 Li Zhang , Yue Li , Huan Zhao , Qing Wang , Lei Xie

For many years, i-vector based audio embedding techniques were the dominant approach for speaker verification and speaker diarization applications. However, mirroring the rise of deep learning in various domains, neural network based audio…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Quan Wang , Carlton Downey , Li Wan , Philip Andrew Mansfield , Ignacio Lopez Moreno

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate…

声音 · 计算机科学 2024-10-08 Ibrahim Aldarmaki , Thamar Solorio , Bhiksha Raj , Hanan Aldarmaki

Deep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly,…

声音 · 计算机科学 2024-07-03 Zhenchun Lei , Hui Yan , Changhong Liu , Yong Zhou , Minglei Ma