中文

VoiceExtender:基于引导扩散模型的短语音文本无关说话人验证

声音 2023-10-10 v1 人工智能 音频与语音处理

摘要

随着 utterance 变短,说话人验证(SV)性能会下降。为此,我们提出一种称为 VoiceExtender 的新架构,为处理短时长语音信号时提升 SV 性能提供了有前景的方案。我们使用两个引导扩散模型——内置及外部说话人嵌入(SE)引导扩散模型,二者均利用基于扩散模型的样本生成器,借助 SE 引导基于短 utterance 增强语音特征。在 VoxCeleb1 数据集上的大量实验结果表明,我们的方法优于基线,在 0.5、1.0、1.5 和 2.0 秒短 utterance 条件下,等错误率(EER)相对分别提升 46.1%、35.7%、10.4% 和 5.7%。

关键词

引用

@article{arxiv.2310.04681,
  title  = {VoiceExtender: Short-utterance Text-independent Speaker Verification with Guided Diffusion Model},
  author = {Yayun He and Zuheng Kang and Jianzong Wang and Junqing Peng and Jing Xiao},
  journal= {arXiv preprint arXiv:2310.04681},
  year   = {2023}
}

备注

Accepted by the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2023)