English

PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components

Audio and Speech Processing 2021-02-17 v1 Machine Learning Sound Signal Processing

Abstract

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by conditioning an acoustic feature. Since a speech waveform contains periodic and aperiodic components, both components should be appropriately modeled to generate a high-quality speech waveform. However, it is difficult to decompose the components from a natural speech waveform in advance. To address this issue, we propose a parallel model and a series model structure separating periodic and aperiodic components. The features of our proposed models are that explicit periodic and aperiodic signals are taken as input, and external periodic/aperiodic decomposition is not needed in training. Experiments using a singing voice corpus show that our proposed structure improves the naturalness of the generated waveform. We also show that the speech waveforms with a pitch outside of the training data range can be generated with more naturalness.

Keywords

Cite

@article{arxiv.2102.07786,
  title  = {PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components},
  author = {Yukiya Hono and Shinji Takaki and Kei Hashimoto and Keiichiro Oura and Yoshihiko Nankaku and Keiichi Tokuda},
  journal= {arXiv preprint arXiv:2102.07786},
  year   = {2021}
}

Comments

5 pages, accepted to ICASSP 2021

R2 v1 2026-06-23T23:11:12.244Z