English

Geneses: Unified Generative Speech Enhancement and Separation

Sound 2026-01-27 v1 Audio and Speech Processing

Abstract

Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches that concatenate speech enhancement (SE) and speech separation (SS) to obtain a clean speech signal for each speaker are promising, conventional SE-SS methods suffer from complex degradations beyond additive noise. To this end, we propose \textbf{Geneses}, a generative framework to achieve unified, high-quality SE--SS. Our Geneses leverages latent flow matching to estimate each speaker's clean speech features using multi-modal diffusion Transformer conditioned on self-supervised learning representation from noisy mixture. We conduct experimental evaluation using two-speaker mixtures from LibriTTS-R under two conditions: additive-noise-only and complex degradations. The results demonstrate that Geneses significantly outperforms a conventional mask-based SE--SS method across various objective metrics with high robustness against complex degradations. Audio samples are available in our demo page.

Keywords

Cite

@article{arxiv.2601.18456,
  title  = {Geneses: Unified Generative Speech Enhancement and Separation},
  author = {Kohei Asai and Wataru Nakata and Yuki Saito and Hiroshi Saruwatari},
  journal= {arXiv preprint arXiv:2601.18456},
  year   = {2026}
}

Comments

Accepted to ICASSP 2025 workshop

R2 v1 2026-07-01T09:20:22.397Z