English

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

Audio and Speech Processing 2025-12-25 v1 Artificial Intelligence Machine Learning

Abstract

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We present GenTSE, a two-stage decoder-only generative LM approach for TSE: Stage-1 predicts coarse semantic tokens, and Stage-2 generates fine acoustic tokens. Separating semantics and acoustics stabilizes decoding and yields more faithful, content-aligned target speech. Both stages use continuous SSL or codec embeddings, offering richer context than discretized-prompt methods. To reduce exposure bias, we employ a Frozen-LM Conditioning training strategy that conditions the LMs on predicted tokens from earlier checkpoints to reduce the gap between teacher-forcing training and autoregressive inference. We further employ DPO to better align outputs with human perceptual preferences. Experiments on Libri2Mix show that GenTSE surpasses previous LM-based systems in speech quality, intelligibility, and speaker consistency.

Keywords

Cite

@article{arxiv.2512.20978,
  title  = {GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model},
  author = {Haoyang Li and Xuyi Zhuang and Azmat Adnan and Ye Ni and Wei Rao and Shreyas Gopal and Eng Siong Chng},
  journal= {arXiv preprint arXiv:2512.20978},
  year   = {2025}
}