English

A Holistic Cascade System, benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation

Computation and Language 2023-01-26 v1 Sound Audio and Speech Processing

Abstract

Expressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy. Existing research in expressive S2ST is limited, typically focusing on a single expressivity aspect at a time. Likewise, this research area lacks standard evaluation protocols and well-curated benchmark datasets. In this work, we propose a holistic cascade system for expressive S2ST, combining multiple prosody transfer techniques previously considered only in isolation. We curate a benchmark expressivity test set in the TV series domain and explored a second dataset in the audiobook domain. Finally, we present a human evaluation protocol to assess multiple expressive dimensions across speech pairs. Experimental results indicate that bi-lingual annotators can assess the quality of expressive preservation in S2ST systems, and the holistic modeling approach outperforms single-aspect systems. Audio samples can be accessed through our demo webpage: https://facebookresearch.github.io/speech_translation/cascade_expressive_s2st.

Keywords

Cite

@article{arxiv.2301.10606,
  title  = {A Holistic Cascade System, benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation},
  author = {Wen-Chin Huang and Benjamin Peloquin and Justine Kao and Changhan Wang and Hongyu Gong and Elizabeth Salesky and Yossi Adi and Ann Lee and Peng-Jen Chen},
  journal= {arXiv preprint arXiv:2301.10606},
  year   = {2023}
}

Comments

This is the full version of our submission to ICASSP 2023

R2 v1 2026-06-28T08:19:55.865Z