English

Read it to me: An emotionally aware Speech Narration Application

Sound 2022-09-08 v1 Computation and Language Machine Learning Audio and Speech Processing

Abstract

In this work we try to perform emotional style transfer on audios. In particular, MelGAN-VC architecture is explored for various emotion-pair transfers. The generated audio is then classified using an LSTM-based emotion classifier for audio. We find that "sad" audio is generated well as compared to "happy" or "anger" as people have similar expressions of sadness.

Keywords

Cite

@article{arxiv.2209.02785,
  title  = {Read it to me: An emotionally aware Speech Narration Application},
  author = {Rishibha Bansal},
  journal= {arXiv preprint arXiv:2209.02785},
  year   = {2022}
}
R2 v1 2026-06-28T00:50:09.758Z