English

Data Augmenting Contrastive Learning of Speech Representations in the Time Domain

Audio and Speech Processing 2020-07-03 v1 Computation and Language Sound

Abstract

Contrastive Predictive Coding (CPC), based on predicting future segments of speech based on past segments is emerging as a powerful algorithm for representation learning of speech signal. However, it still under-performs other methods on unsupervised evaluation benchmarks. Here, we introduce WavAugment, a time-domain data augmentation library and find that applying augmentation in the past is generally more efficient and yields better performances than other methods. We find that a combination of pitch modification, additive noise and reverberation substantially increase the performance of CPC (relative improvement of 18-22%), beating the reference Libri-light results with 600 times less data. Using an out-of-domain dataset, time-domain data augmentation can push CPC to be on par with the state of the art on the Zero Speech Benchmark 2017. We also show that time-domain data augmentation consistently improves downstream limited-supervision phoneme classification tasks by a factor of 12-15% relative.

Keywords

Cite

@article{arxiv.2007.00991,
  title  = {Data Augmenting Contrastive Learning of Speech Representations in the Time Domain},
  author = {Eugene Kharitonov and Morgane Rivière and Gabriel Synnaeve and Lior Wolf and Pierre-Emmanuel Mazaré and Matthijs Douze and Emmanuel Dupoux},
  journal= {arXiv preprint arXiv:2007.00991},
  year   = {2020}
}
R2 v1 2026-06-23T16:47:44.298Z