English

Exploring Efficient Directional and Distance Cues for Regional Speech Separation

Sound 2025-08-12 v1 Audio and Speech Processing

Abstract

In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound source not only from specified direction but also within defined distance. Specifically, our method employs an improved delay-and-sum technique to obtain directional cues, substantially enhancing the signal from the target direction. We further enhance separation by incorporating the direct-to-reverberant ratio into the input features, enabling the model to better discriminate sources within and beyond a specified distance. Experimental results demonstrate that our proposed method leads to substantial gains across multiple objective metrics. Furthermore, our method achieves state-of-the-art performance on the CHiME-8 MMCSG dataset, which was recorded in real-world conversational scenarios, underscoring its effectiveness for speech separation in practical applications.

Keywords

Cite

@article{arxiv.2508.07563,
  title  = {Exploring Efficient Directional and Distance Cues for Regional Speech Separation},
  author = {Yiheng Jiang and Haoxu Wang and Yafeng Chen and Gang Qiao and Biao Tian},
  journal= {arXiv preprint arXiv:2508.07563},
  year   = {2025}
}

Comments

This paper has been accepted by Interspeech 2025

R2 v1 2026-07-01T04:43:31.937Z