Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling
Abstract
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of- candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.
Cite
@article{arxiv.2607.06088,
title = {Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling},
author = {Anastasia Zorkina and Alexandr Anikin and Nikita Khmelev and Anastasiya Korenevskaya and Sergey Novoselov and Vladimir Volokhov and Maxim Korenevsky and Yuriy Matveev},
journal= {arXiv preprint arXiv:2607.06088},
year = {2026}
}
Comments
Accepted at the ICML 2026 Workshop on Machine Learning for Audio