Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
Abstract
Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ability of such systems remains underexplored, primarily due to the lack of explicit speaker-specific supervision in the FM framework. To this end, we conduct an empirical analysis of speaker information distribution and reveal its non-uniform allocation across time steps and network layers, underscoring the need for adaptive speaker alignment. Accordingly, we propose Time-Layer Adaptive Speaker Alignment (TLA-SA), a strategy that enhances speaker consistency by jointly leveraging temporal and hierarchical variations. Experimental results show that TLA-SA substantially improves speaker similarity over baseline systems on both research- and industrial-scale datasets and generalizes well across diverse model architectures, including decoder-only language model (LM)-based and free TTS systems. A demo is provided.
Keywords
Cite
@article{arxiv.2511.09995,
title = {Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS},
author = {Haoyu Li and Mingyang Han and Yu Xi and Dongxiao Wang and Hankun Wang and Haoxiang Shi and Boyu Li and Jun Song and Bo Zheng and Shuai Wang and Kai Yu},
journal= {arXiv preprint arXiv:2511.09995},
year = {2026}
}
Comments
Submitted to INTERSPEECH 2026