用于说话人识别的判别训练 i-vector 提取器的因式分解
音频与语音处理
2019-04-10 v1 声音
摘要
在这项工作中,我们延续对用于说话人验证(SV)的 i-vector 提取器的研究,并优化其架构以实现快速而有效的判别训练。我们的动机源于原始生成式 i-vector 模型大量参数所带来的计算与内存需求。我们的目标是保留原始生成模型的性能,同时使模型聚焦于提取与说话人相关的信息。我们表明,可以用参数显著更少的模型来表示标准生成式 i-vector 提取器,并在 SV 任务上获得相似性能。我们可通过判别训练进一步精炼这一紧凑模型,并获得能在代表不同声学领域的各类 SV 基准上带来更优性能的 i-vector。
引用
@article{arxiv.1904.04235,
title = {Factorization of Discriminatively Trained i-vector Extractor for Speaker Recognition},
author = {Ondrej Novotny and Oldrich Plchot and Ondrej Glembek and Lukas Burget},
journal= {arXiv preprint arXiv:1904.04235},
year = {2019}
}
备注
Submitted to Interspeech 2019, Graz, Austria. arXiv admin note: substantial text overlap with arXiv:1810.13183