English

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

Sound 2026-05-19 v10 Audio and Speech Processing

Abstract

This paper presents the Multi-Language Audio Anti-Spoofing Dataset (MLAAD), version 10: a dataset of synthetic audio to train and evaluate audio deepfake detection models. It features 175 Text-to-Speech (TTS) models, comprising a total of 1002.9 hours of synthetic voice in 54 different languages. To evaluate this dataset, we train three state-of-the-art deepfake detection models with MLAAD and observe that it demonstrates superior performance to comparable datasets like InTheWild and FakeOrReal when used as a training resource. Moreover, compared to the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, each excelling on four datasets. By publishing the dataset and making a trained model accessible via an interactive webserver, we aim to democratize anti-spoofing technology, making it accessible beyond the realm of specialists, and contributing to global efforts against audio spoofing and deepfakes.

Keywords

Cite

@article{arxiv.2401.09512,
  title  = {MLAAD: The Multi-Language Audio Anti-Spoofing Dataset},
  author = {Nicolas M. Müller and Piotr Kawa and Wei Herng Choong and Edresson Casanova and Eren Gölge and Thorsten Müller and Piotr Syga and Philip Sperl and Konstantin Böttinger},
  journal= {arXiv preprint arXiv:2401.09512},
  year   = {2026}
}

Comments

IJCNN 2024

R2 v1 2026-06-28T14:19:43.192Z