English

CAMP: a Two-Stage Approach to Modelling Prosody in Context

Audio and Speech Processing 2021-02-15 v2

Abstract

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background noise); (2) determining appropriate prosody without sufficient context is an ill-posed problem. In this paper, we propose solutions to both these issues. To mitigate the challenge of modelling a slow-varying signal, we learn to disentangle prosodic information using a word level representation. To alleviate the ill-posed nature of prosody modelling, we use syntactic and semantic information derived from text to learn a context-dependent prior over our prosodic space. Our Context-Aware Model of Prosody (CAMP) outperforms the state-of-the-art technique, closing the gap with natural speech by 26%. We also find that replacing attention with a jointly-trained duration model improves prosody significantly.

Keywords

Cite

@article{arxiv.2011.01175,
  title  = {CAMP: a Two-Stage Approach to Modelling Prosody in Context},
  author = {Zack Hodari and Alexis Moinet and Sri Karlapati and Jaime Lorenzo-Trueba and Thomas Merritt and Arnaud Joly and Ammar Abbas and Penny Karanasou and Thomas Drugman},
  journal= {arXiv preprint arXiv:2011.01175},
  year   = {2021}
}

Comments

5 pages. Published in the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2021)

R2 v1 2026-06-23T19:51:29.290Z