English

Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet

Computation and Language 2025-02-06 v1 Artificial Intelligence Machine Learning Sound Audio and Speech Processing

Abstract

We present lightweight flow matching multilingual text-to-speech (TTS) systems for Ojibwe, Mi'kmaq, and Maliseet, three Indigenous languages in North America. Our results show that training a multilingual TTS model on three typologically similar languages can improve the performance over monolingual models, especially when data are scarce. Attention-free architectures are highly competitive with self-attention architecture with higher memory efficiency. Our research not only advances technical development for the revitalization of low-resource languages but also highlights the cultural gap in human evaluation protocols, calling for a more community-centered approach to human evaluation.

Keywords

Cite

@article{arxiv.2502.02703,
  title  = {Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet},
  author = {Shenran Wang and Changbing Yang and Mike Parkhill and Chad Quinn and Christopher Hammerly and Jian Zhu},
  journal= {arXiv preprint arXiv:2502.02703},
  year   = {2025}
}
R2 v1 2026-06-28T21:32:43.304Z