English

Many-Speakers Single Channel Speech Separation with Optimal Permutation Training

Sound 2021-11-09 v4 Artificial Intelligence Machine Learning Audio and Speech Processing

Abstract

Single channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the current methods, which rely on the Permutation Invariant Loss (PIT). In this work, we present a permutation invariant training that employs the Hungarian algorithm in order to train with an O(C3)O(C^3) time complexity, where CC is the number of speakers, in comparison to O(C!)O(C!) of PIT based methods. Furthermore, we present a modified architecture that can handle the increased number of speakers. Our approach separates up to 2020 speakers and improves the previous results for large CC by a wide margin.

Keywords

Cite

@article{arxiv.2104.08955,
  title  = {Many-Speakers Single Channel Speech Separation with Optimal Permutation Training},
  author = {Shaked Dovrat and Eliya Nachmani and Lior Wolf},
  journal= {arXiv preprint arXiv:2104.08955},
  year   = {2021}
}

Comments

Accepted to Interspeech 2021, Data creation link added

R2 v1 2026-06-24T01:18:15.476Z