English

M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis

Sound 2025-12-05 v1

Abstract

Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing methods depend on duration modeling or pseudo-alignment strategies that severely limit naturalness and computational efficiency. We propose M3-TTS, a concise and efficient NAR TTS paradigm based on multi-modal diffusion transformer (MM-DiT) architecture. M3-TTS employs joint diffusion transformer layers for cross-modal alignment, achieving stable monotonic alignment between variable-length text-speech sequences without pseudo-alignment requirements. Single diffusion transformer layers further enhance acoustic detail modeling. The framework integrates a mel-vae codec that provides 3* training acceleration. Experimental results on Seed-TTS and AISHELL-3 benchmarks demonstrate that M3-TTS achieves state-of-the-art NAR performance with the lowest word error rates (1.36\% English, 1.31\% Chinese) while maintaining competitive naturalness scores. Code and demos will be available at https://wwwwxp.github.io/M3-TTS.

Keywords

Cite

@article{arxiv.2512.04720,
  title  = {M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis},
  author = {Xiaopeng Wang and Chunyu Qiang and Ruibo Fu and Zhengqi Wen and Xuefei Liu and Yukun Liu and Yuzhe Liang and Kang Yin and Yuankun Xie and Heng Xie and Chenxing Li and Chen Zhang and Changsheng Li},
  journal= {arXiv preprint arXiv:2512.04720},
  year   = {2025}
}

Comments

Submitted to ICASSP 2026