English

Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer

Audio and Speech Processing 2023-06-22 v2

Abstract

Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing and transferring details and variations in prosody is a challenge. We introduce phrase-level cross-lingual prosody transfer for expressive multi-lingual machine dubbing. The proposed phrase-level prosody transfer delivers a significant 6.2% MUSHRA score increase over a baseline with utterance-level global prosody transfer, thereby closing the gap between the baseline and expressive human dubbing by 23.2%, while preserving intelligibility of the synthesised speech.

Keywords

Cite

@article{arxiv.2306.11662,
  title  = {Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer},
  author = {Jakub Swiatkowski and Duo Wang and Mikolaj Babianski and Giuseppe Coccia and Patrick Lumban Tobing and Ravichander Vipperla and Viacheslav Klimkov and Vincent Pollet},
  journal= {arXiv preprint arXiv:2306.11662},
  year   = {2023}
}

Comments

Accepted to INTERSPEECH 2023

R2 v1 2026-06-28T11:09:50.866Z