中文

Speech-XLNet:用于自注意力网络的无监督声学模型预训练

计算与语言 2020-05-25 v2 音频与语音处理

摘要

自注意力网络(SAN)可通过 BERT 和 XLNet 等无监督预训练范式的双向表示学习获得显著收益。在本文中,我们提出一种类 XLNet 的预训练方案“Speech-XLNet”,用于无监督声学模型预训练,以通过 SAN 学习语音表示。预训练后的 SAN 在混合 SAN/HMM 框架下微调。我们推测,通过打乱语音帧顺序,Speech-XLNet 中的排列作为一种强正则化器,促使 SAN 通过其注意力权重关注全局结构来进行推断。此外,Speech-XLNet 还允许模型探索双向上下文以进行有效语音表示学习。在 TIMIT 和 WSJ 上的实验表明,与从随机初始化权重训练的系统相比,Speech-XLNet 在收敛速度与识别准确率两方面均大幅提升了 SAN/HMM 性能。我们的最佳系统在 TIMIT 和 WSJ 任务上分别实现了 11.9% 和 8.3% 的相对提升。特别地,最佳系统在 TIMIT 测试集上取得了 13.3% 的音素错误率(PER),据我们所知,这是单一系统所获得的最低 PER。

关键词

引用

@article{arxiv.1910.10387,
  title  = {Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks},
  author = {Xingchen Song and Guangsen Wang and Zhiyong Wu and Yiheng Huang and Dan Su and Dong Yu and Helen Meng},
  journal= {arXiv preprint arXiv:1910.10387},
  year   = {2020}
}

备注

\c{opyright} 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works