English

VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation

Sound 2023-03-15 v1 Machine Learning Audio and Speech Processing

Abstract

We introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM and supports explicit control of accent, language, speaker and fine-grained F0F_0 and energy features for speech synthesis. We utilize the Indic languages dataset, released for LIMMITS 2023 as part of ICASSP Signal Processing Grand Challenge, to synthesize speech in 3 different languages. Our model supports transferring the language of a speaker while retaining their voice and the native accent of the target language. We utilize the large-parameter RADMMM model for Track 11 and lightweight VANI model for Track 22 and 33 of the competition.

Keywords

Cite

@article{arxiv.2303.07578,
  title  = {VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation},
  author = {Rohan Badlani and Akshit Arora and Subhankar Ghosh and Rafael Valle and Kevin J. Shih and João Felipe Santos and Boris Ginsburg and Bryan Catanzaro},
  journal= {arXiv preprint arXiv:2303.07578},
  year   = {2023}
}

Comments

Presentation accepted at ICASSP 2023