中文

HighRateMOS:面向语音质量评估的采样率感知建模

音频与语音处理 2025-06-30 v1

摘要

现代语音质量预测模型在训练时使用特定采样率的音频数据。在测试时面对更高采样率的音频,这些模型会产生偏差评分。我们提出了HighRateMOS,是首个显式考虑采样率的非侵入式mean opinion score(MOS)模型。HighRateMOS集成了三个模型变体,利用以下信息:(i)语音采样率的可学习嵌入,(ii)Wav2vec 2.0自监督嵌入,(iii)多尺度CNN频谱特征,以及(iv)MFCC特征。在AudioMOS 2025 Track3中,HighRateMOS在五个八个指标中名列前茅。我们的实验确认,直接建模采样率导致更稳健、对采样率agnostic的语音质量预测。

关键词

引用

@article{arxiv.2506.21951,
  title  = {HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment},
  author = {Wenze Ren and Yi-Cheng Lin and Wen-Chin Huang and Ryandhimas E. Zezario and Szu-Wei Fu and Sung-Feng Huang and Erica Cooper and Haibin Wu and Hung-Yu Wei and Hsin-Min Wang and Hung-yi Lee and Yu Tsao},
  journal= {arXiv preprint arXiv:2506.21951},
  year   = {2025}
}

备注

Under Review, 3 pages + 1 References