English

Detection of AI-Synthesized Speech Using Cepstral & Bispectral Statistics

Machine Learning 2021-04-13 v2 Multimedia Audio and Speech Processing Machine Learning

Abstract

Digital technology has made possible unimaginable applications come true. It seems exciting to have a handful of tools for easy editing and manipulation, but it raises alarming concerns that can propagate as speech clones, duplicates, or maybe deep fakes. Validating the authenticity of a speech is one of the primary problems of digital audio forensics. We propose an approach to distinguish human speech from AI synthesized speech exploiting the Bi-spectral and Cepstral analysis. Higher-order statistics have less correlation for human speech in comparison to a synthesized speech. Also, Cepstral analysis revealed a durable power component in human speech that is missing for a synthesized speech. We integrate both these analyses and propose a machine learning model to detect AI synthesized speech.

Keywords

Cite

@article{arxiv.2009.01934,
  title  = {Detection of AI-Synthesized Speech Using Cepstral & Bispectral Statistics},
  author = {Arun Kumar Singh and Priyanka Singh},
  journal= {arXiv preprint arXiv:2009.01934},
  year   = {2021}
}

Comments

6 Pages, 6 Figures, 1 Table