English

CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

Computation and Language 2024-06-19 v2 Sound Audio and Speech Processing

Abstract

This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.

Keywords

Cite

@article{arxiv.2406.00021,
  title  = {CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning},
  author = {Medha Hira and Arnav Goel and Anubha Gupta},
  journal= {arXiv preprint arXiv:2406.00021},
  year   = {2024}
}

Comments

8 pages, Accepted at ICLR 2024 - Tiny Track

R2 v1 2026-06-28T16:48:52.780Z