中文

增强型生成式机器听者

音频与语音处理 2026-01-26 v2 人工智能 机器学习

摘要

我们提出了 GMLv2,这是一个基于参考的模型,用于预测主观音频质量,即由 MUSHRA 分数衡量。GMLv2 引入基于 Beta 分布的损失函数以建模听众评分,并整合额外的神经音频编码(NAC)主观数据集,以扩展其泛化能力和适用性。在多样化测试集上进行的广泛评估表明,提出的 GMLv2 在与主观分数相关性以及在各种内容类型和编解码器配置下可靠预测这些分数的方面, consistently 优于 PEAQ 和 ViSQOL 等常用指标。因此,GMLv2 提供了一个可扩展且自动化的感知音频质量评估框架,已被证实可加速现代音频编码技术的研究与开发。

关键词

引用

@article{arxiv.2509.21463,
  title  = {Enhanced Generative Machine Listener},
  author = {Vishnu Raj and Gouthaman KV and Shiv Gehlot and Lars Villemoes and Arijit Biswas},
  journal= {arXiv preprint arXiv:2509.21463},
  year   = {2026}
}

备注

Accepted to the 51st IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 4-8 May 2026