English

Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS

Sound 2026-04-03 v1 Audio and Speech Processing

Abstract

We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by SigLIP-style cross-modal contrastive learning using mixed-phoneme batches, with an additional same-phoneme refinement stage studied separately. We evaluate intrinsic text-audio retrieval and downstream synthesis in Grad-TTS and a latent diffusion TTS system. The two-stage curriculum (MLM + mixed-phoneme contrastive learning) achieves the best overall synthesis quality in terms of intelligibility, speaker similarity, and perceptual measures. Although same-phoneme refinement improves prosodic retrieval, it reduces phoneme discrimination and degrades synthesis. These findings indicate that improvements in embedding-space metrics do not necessarily translate to better generative performance and highlight the need to balance phoneme discrimination and prosodic sensitivity in TTS pretraining.

Keywords

Cite

@article{arxiv.2604.01247,
  title  = {Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS},
  author = {Kirill Borodin and Vasiliy Kudryavtsev and Maxim Maslov and Nikita Vasiliev and Mikhail Gorodnichev and Grach Mkrtchian},
  journal= {arXiv preprint arXiv:2604.01247},
  year   = {2026}
}

Comments

This paper has been submitted to Interspeech 2026 for review

R2 v1 2026-07-01T11:49:33.893Z