English

Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition

Audio and Speech Processing 2022-04-11 v2 Machine Learning Sound

Abstract

Speech enhancement (SE) aims to suppress the additive noise from a noisy speech signal to improve the speech's perceptual quality and intelligibility. However, the over-suppression phenomenon in the enhanced speech might degrade the performance of downstream automatic speech recognition (ASR) task due to the missing latent information. To alleviate such problem, we propose an interactive feature fusion network (IFF-Net) for noise-robust speech recognition to learn complementary information from the enhanced feature and original noisy feature. Experimental results show that the proposed method achieves absolute word error rate (WER) reduction of 4.1% over the best baseline on RATS Channel-A corpus. Our further analysis indicates that the proposed IFF-Net can complement some missing information in the over-suppressed enhanced feature.

Keywords

Cite

@article{arxiv.2110.05267,
  title  = {Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition},
  author = {Yuchen Hu and Nana Hou and Chen Chen and Eng Siong Chng},
  journal= {arXiv preprint arXiv:2110.05267},
  year   = {2022}
}

Comments

5 pages, 7 figures, Accepted by ICASSP 2022

R2 v1 2026-06-24T06:47:35.303Z