English

Ranking the Impact of Contextual Specialization in Neural Speech Enhancement

Audio and Speech Processing 2026-07-06 v1 Sound

Abstract

We systematically investigate neural speech enhancement systems, ranging from very small (\sim10\,k parameters) to medium-large (\sim2-5\,M parameters), which specialize to acoustic conditions using contextual information such as speaker identity, noise type, speaker gender, spoken language, and SNR. By fine-tuning generalist models on specific data subsets, we find that specializing to a speaker's identity consistently yields the largest gains in estimated speech intelligibility and quality. In contrast, specializing to SNR, noise type, or gender offers only marginal benefits. Crucially, we show that a small model specialized to both a specific speaker and a specific noise type can match or exceed the performance of a generalist model ten times its size. Further, cross-lingual tests reveal that models specialized to a target language outperform multilingual generalists, suggesting that language is a salient feature for specialization. These findings highlight the potential of small, adaptive models for resource-constrained applications like hearing aids, which specialize on-the-fly to contextual information.

Keywords

Cite

@article{arxiv.2607.04826,
  title  = {Ranking the Impact of Contextual Specialization in Neural Speech Enhancement},
  author = {Peter Leer and Svend Feldt and Zheng-Hua Tan and Jan Østergaard and Jesper Jensen},
  journal= {arXiv preprint arXiv:2607.04826},
  year   = {2026}
}

Comments

Accepted to ICASSP 2026