English

Efficient acoustic feature transformation in mismatched environments using a Guided-GAN

Sound 2022-10-07 v3 Machine Learning Audio and Speech Processing

Abstract

We propose a new framework to improve automatic speech recognition (ASR) systems in resource-scarce environments using a generative adversarial network (GAN) operating on acoustic input features. The GAN is used to enhance the features of mismatched data prior to decoding, or can optionally be used to fine-tune the acoustic model. We achieve improvements that are comparable to multi-style training (MTR), but at a lower computational cost. With less than one hour of data, an ASR system trained on good quality data, and evaluated on mismatched audio is improved by between 11.5% and 19.7% relative word error rate (WER). Experiments demonstrate that the framework can be very useful in under-resourced environments where training data and computational resources are limited. The GAN does not require parallel training data, because it utilises a baseline acoustic model to provide an additional loss term that guides the generator to create acoustic features that are better classified by the baseline.

Keywords

Cite

@article{arxiv.2210.00721,
  title  = {Efficient acoustic feature transformation in mismatched environments using a Guided-GAN},
  author = {Walter Heymans and Marelie H. Davel and Charl van Heerden},
  journal= {arXiv preprint arXiv:2210.00721},
  year   = {2022}
}

Comments

Final published version available at: Efficient acoustic feature transformation in mismatched environments using a Guided-GAN. Speech Communication, 143, pp.10-20

R2 v1 2026-06-28T02:34:49.816Z