English

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models

Sound 2026-05-01 v1 Machine Learning Audio and Speech Processing

Abstract

To train transcriptor models that produce robust results, a large and diverse labeled dataset is required. Finding such data with the necessary characteristics is a challenging task, especially for languages less popular than English. Moreover, producing such data requires significant effort and often money. Therefore, a strategy to mitigate this problem is the use of data augmentation techniques. In this work, we propose a framework that approaches data augmentation based on deepfake audio. To validate the produced framework, experiments were conducted using existing deepfake and transcription models. A voice cloner and a dataset produced by Indians (in English) were selected, ensuring the presence of a single accent in the dataset. Subsequently, the augmented data was used to train speech to text models in various scenarios.

Keywords

Cite

@article{arxiv.2309.12802,
  title  = {Deepfake audio as a data augmentation technique for training automatic speech to text transcription models},
  author = {Alexandre R. Ferreira and Cláudio E. C. Campelo},
  journal= {arXiv preprint arXiv:2309.12802},
  year   = {2026}
}

Comments

9 pages, 6 figures, 7 tables

R2 v1 2026-06-28T12:29:21.621Z