HighRateMOS:面向语音质量评估的采样率感知建模
音频与语音处理
2025-06-30 v1
摘要
现代语音质量预测模型在训练时使用特定采样率的音频数据。在测试时面对更高采样率的音频,这些模型会产生偏差评分。我们提出了HighRateMOS,是首个显式考虑采样率的非侵入式mean opinion score(MOS)模型。HighRateMOS集成了三个模型变体,利用以下信息:(i)语音采样率的可学习嵌入,(ii)Wav2vec 2.0自监督嵌入,(iii)多尺度CNN频谱特征,以及(iv)MFCC特征。在AudioMOS 2025 Track3中,HighRateMOS在五个八个指标中名列前茅。我们的实验确认,直接建模采样率导致更稳健、对采样率agnostic的语音质量预测。
关键词
引用
@article{arxiv.2506.21951,
title = {HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment},
author = {Wenze Ren and Yi-Cheng Lin and Wen-Chin Huang and Ryandhimas E. Zezario and Szu-Wei Fu and Sung-Feng Huang and Erica Cooper and Haibin Wu and Hung-Yu Wei and Hsin-Min Wang and Hung-yi Lee and Yu Tsao},
journal= {arXiv preprint arXiv:2506.21951},
year = {2025}
}
备注
Under Review, 3 pages + 1 References