English

Towards end-to-end F0 voice conversion based on Dual-GAN with convolutional wavelet kernels

Audio and Speech Processing 2021-04-16 v1 Machine Learning Sound

Abstract

This paper presents a end-to-end framework for the F0 transformation in the context of expressive voice conversion. A single neural network is proposed, in which a first module is used to learn F0 representation over different temporal scales and a second adversarial module is used to learn the transformation from one emotion to another. The first module is composed of a convolution layer with wavelet kernels so that the various temporal scales of F0 variations can be efficiently encoded. The single decomposition/transformation network allows to learn in a end-to-end manner the F0 decomposition that are optimal with respect to the transformation, directly from the raw F0 signal.

Keywords

Cite

@article{arxiv.2104.07283,
  title  = {Towards end-to-end F0 voice conversion based on Dual-GAN with convolutional wavelet kernels},
  author = {Clément Le Moine Veillon and Nicolas Obin and Axel Roebel},
  journal= {arXiv preprint arXiv:2104.07283},
  year   = {2021}
}
R2 v1 2026-06-24T01:11:22.047Z