English

MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction

Audio and Speech Processing 2024-01-26 v2 Multimedia

Abstract

Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to assess how close synthesized speech is to the natural human voice. We propose MOS-FAD, where MOS can be leveraged at two key points in FAD: training data selection and model fusion. In training data selection, we demonstrate that MOS enables effective filtering of samples from unbalanced datasets. In the model fusion, our results demonstrate that incorporating MOS as a gating mechanism in FAD model fusion enhances overall performance.

Keywords

Cite

@article{arxiv.2401.13249,
  title  = {MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction},
  author = {Wangjin Zhou and Zhengdong Yang and Chenhui Chu and Sheng Li and Raj Dabre and Yi Zhao and Tatsuya Kawahara},
  journal= {arXiv preprint arXiv:2401.13249},
  year   = {2024}
}

Comments

Accepted in ICASSP2024