English

TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

Sound 2024-09-18 v3 Machine Learning Audio and Speech Processing

Abstract

We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.

Keywords

Cite

@article{arxiv.2409.07841,
  title  = {TSELM: Target Speaker Extraction using Discrete Tokens and Language Models},
  author = {Beilong Tang and Bang Zeng and Ming Li},
  journal= {arXiv preprint arXiv:2409.07841},
  year   = {2024}
}

Comments

Submitted to ICASSP 2025

R2 v1 2026-06-28T18:42:10.728Z