English

Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation

Sound 2021-09-10 v3 Computation and Language Audio and Speech Processing

Abstract

In this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher model. The teacher model then estimates separated sources that are used to train a student model with standard permutation invariant training (PIT). The student model can be fine-tuned with supervised data, i.e., paired artificial mixtures and clean speech sources, and further improved via model distillation. Experiments with single and multi channel mixtures show that the teacher-student training resolves the over-separation problem observed in the original MixIT method. Further, the semisupervised performance is comparable to a fully-supervised separation system trained using ten times the amount of supervised data.

Keywords

Cite

@article{arxiv.2106.07843,
  title  = {Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation},
  author = {Jisi Zhang and Catalin Zorila and Rama Doddipatla and Jon Barker},
  journal= {arXiv preprint arXiv:2106.07843},
  year   = {2021}
}

Comments

Accepted to Interspeech 2021

R2 v1 2026-06-24T03:12:13.547Z