MOS 预测网络的泛化能力
音频与语音处理
2022-02-15 v3
摘要
预测合成语音听者意见的自动方法仍难以实现,因为听者、被评测系统、语音特征,甚至所给指令与评分量表在每次测试中均有所不同。尽管诸如平均意见分(MOS)等指标的自动预测器在同一测试的样片上可达到高预测准确率,它们通常不能很好泛化到新的听音测试情境。本文使用多种用于 MOS 预测的神经网络(包括 MOSNet 与诸如 wav2vec2 的自监督语音模型),在零样本与微调设定下考察它们在不同听音测试数据上的表现。我们发现,即便在句级预测的零样本这一最具挑战情形下,为 MOS 预测微调的 wav2vec2 模型对域外数据也具有良好泛化能力,且对域内数据微调可改善预测。我们还观察到,未见系统对 MOS 预测模型尤其具有挑战性。
引用
@article{arxiv.2110.02635,
title = {Generalization Ability of MOS Prediction Networks},
author = {Erica Cooper and Wen-Chin Huang and Tomoki Toda and Junichi Yamagishi},
journal= {arXiv preprint arXiv:2110.02635},
year = {2022}
}
备注
\c{opyright} 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works