SpMis: An Investigation of Synthetic Spoken Misinformation Detection
Abstract
In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, they have also raised concerns about the misuse of this technology, particularly for generating synthetic misinformation. Current research primarily focuses on distinguishing machine-generated speech from human-produced speech, but the more urgent challenge is detecting misinformation within spoken content. This task requires a thorough analysis of factors such as speaker identity, topic, and synthesis. To address this need, we conduct an initial investigation into synthetic spoken misinformation detection by introducing an open-source dataset, SpMis. SpMis includes speech synthesized from over 1,000 speakers across five common topics, utilizing state-of-the-art text-to-speech systems. Although our results show promising detection capabilities, they also reveal substantial challenges for practical implementation, underscoring the importance of ongoing research in this critical area.
Keywords
Cite
@article{arxiv.2409.11308,
title = {SpMis: An Investigation of Synthetic Spoken Misinformation Detection},
author = {Peizhuo Liu and Li Wang and Renqiang He and Haorui He and Lei Wang and Huadi Zheng and Jie Shi and Tong Xiao and Zhizheng Wu},
journal= {arXiv preprint arXiv:2409.11308},
year = {2024}
}
Comments
Accepted in SLT 2024