English

Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

Sound 2026-07-07 v1

Abstract

Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-NN candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.

Cite

@article{arxiv.2607.06088,
  title  = {Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling},
  author = {Anastasia Zorkina and Alexandr Anikin and Nikita Khmelev and Anastasiya Korenevskaya and Sergey Novoselov and Vladimir Volokhov and Maxim Korenevsky and Yuriy Matveev},
  journal= {arXiv preprint arXiv:2607.06088},
  year   = {2026}
}

Comments

Accepted at the ICML 2026 Workshop on Machine Learning for Audio