Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
Abstract
Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS models--Dia2, Maya1, and MeloTTS--representing streaming, LLM-based, and non-autoregressive architectures. A corpus of 12,000 synthetic audio samples was generated using the Daily-Dialog dataset and evaluated against four detection frameworks, including semantic, structural, and signal-level approaches. The results reveal significant variability in detector performance across generative mechanisms: models effective against one TTS architecture may fail against others, particularly LLM-based synthesis. In contrast, a multi-view detection approach combining complementary analysis levels demonstrates robust performance across all evaluated models. These findings highlight the limitations of single-paradigm detectors and emphasize the necessity of integrated detection strategies to address the evolving landscape of audio deepfake threats.
Cite
@article{arxiv.2601.20510,
title = {Audio Deepfake Detection in the Age of Advanced Text-to-Speech models},
author = {Robin Singh and Aditya Yogesh Nair and Fabio Palumbo and Florian Barbaro and Anna Dyka and Lohith Rachakonda},
journal= {arXiv preprint arXiv:2601.20510},
year = {2026}
}
Comments
This work was performed using HPC resources from GENCI-IDRIS (Grant 2025- AD011016076)