English

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Audio and Speech Processing 2024-09-25 v1 Sound

Abstract

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at \url{https://github.com/wenet-e2e/WeSep.}

Keywords

Cite

@article{arxiv.2409.15799,
  title  = {WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction},
  author = {Shuai Wang and Ke Zhang and Shaoxiong Lin and Junjie Li and Xuefei Wang and Meng Ge and Jianwei Yu and Yanmin Qian and Haizhou Li},
  journal= {arXiv preprint arXiv:2409.15799},
  year   = {2024}
}

Comments

Interspeech 2024

R2 v1 2026-06-28T18:54:54.198Z