MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
Abstract
This paper presents the Multi-Language Audio Anti-Spoofing Dataset (MLAAD), version 10: a dataset of synthetic audio to train and evaluate audio deepfake detection models. It features 175 Text-to-Speech (TTS) models, comprising a total of 1002.9 hours of synthetic voice in 54 different languages. To evaluate this dataset, we train three state-of-the-art deepfake detection models with MLAAD and observe that it demonstrates superior performance to comparable datasets like InTheWild and FakeOrReal when used as a training resource. Moreover, compared to the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, each excelling on four datasets. By publishing the dataset and making a trained model accessible via an interactive webserver, we aim to democratize anti-spoofing technology, making it accessible beyond the realm of specialists, and contributing to global efforts against audio spoofing and deepfakes.
Keywords
Cite
@article{arxiv.2401.09512,
title = {MLAAD: The Multi-Language Audio Anti-Spoofing Dataset},
author = {Nicolas M. Müller and Piotr Kawa and Wei Herng Choong and Edresson Casanova and Eren Gölge and Thorsten Müller and Piotr Syga and Philip Sperl and Konstantin Böttinger},
journal= {arXiv preprint arXiv:2401.09512},
year = {2026}
}
Comments
IJCNN 2024