English

Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS

Audio and Speech Processing 2026-03-18 v3

Abstract

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ability of such systems remains underexplored, primarily due to the lack of explicit speaker-specific supervision in the FM framework. To this end, we conduct an empirical analysis of speaker information distribution and reveal its non-uniform allocation across time steps and network layers, underscoring the need for adaptive speaker alignment. Accordingly, we propose Time-Layer Adaptive Speaker Alignment (TLA-SA), a strategy that enhances speaker consistency by jointly leveraging temporal and hierarchical variations. Experimental results show that TLA-SA substantially improves speaker similarity over baseline systems on both research- and industrial-scale datasets and generalizes well across diverse model architectures, including decoder-only language model (LM)-based and free TTS systems. A demo is provided.

Keywords

Cite

@article{arxiv.2511.09995,
  title  = {Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS},
  author = {Haoyu Li and Mingyang Han and Yu Xi and Dongxiao Wang and Hankun Wang and Haoxiang Shi and Boyu Li and Jun Song and Bo Zheng and Shuai Wang and Kai Yu},
  journal= {arXiv preprint arXiv:2511.09995},
  year   = {2026}
}

Comments

Submitted to INTERSPEECH 2026

R2 v1 2026-07-01T07:35:08.717Z