中文

论低资源语音识别中对比表示的缩放

音频与语音处理 2021-02-02 v1 机器学习 声音

摘要

近期通过对比训练的自监督学习进展表明,仅用 10 分钟标注数据即可学习到具有竞争力的语音识别系统。然而,这些系统计算开销大,因为它们需要在大参数空间中进行预训练而后微调。我们在不进行微调的情况下探索此类系统的性能,即在计算密集的 wav2vec 2.0 框架的固定表示上训练最先进的语音识别器。我们发现不加微调时性能下降,且在极低资源设定下 wav2vec 2.0 不如其前代。此外,我们发现 wav2vec 2.0 表示存在于低维子空间中,且对表示特征去相关可稳定自动语音识别器的训练。最后,我们提出对原始 wav2vec 框架的双向扩展,可一致地改善性能。

关键词

引用

@article{arxiv.2102.00850,
  title  = {On Scaling Contrastive Representations for Low-Resource Speech Recognition},
  author = {Lasse Borgholt and Tycho Max Sylvester Tax and Jakob Drachmann Havtorn and Lars Maaløe and Christian Igel},
  journal= {arXiv preprint arXiv:2102.00850},
  year   = {2021}
}

备注

{\copyright} 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works