English

Deep Speech Synthesis from Multimodal Articulatory Representations

Audio and Speech Processing 2024-12-19 v1 Sound

Abstract

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics.

Keywords

Cite

@article{arxiv.2412.13387,
  title  = {Deep Speech Synthesis from Multimodal Articulatory Representations},
  author = {Peter Wu and Bohan Yu and Kevin Scheck and Alan W Black and Aditi S. Krishnapriyan and Irene Y. Chen and Tanja Schultz and Shinji Watanabe and Gopala K. Anumanchipalli},
  journal= {arXiv preprint arXiv:2412.13387},
  year   = {2024}
}
R2 v1 2026-06-28T20:39:38.870Z