English

Spiking Structured State Space Model for Monaural Speech Enhancement

Sound 2024-04-23 v2 Computer Vision and Pattern Recognition Audio and Speech Processing

Abstract

Speech enhancement seeks to extract clean speech from noisy signals. Traditional deep learning methods face two challenges: efficiently using information in long speech sequences and high computational costs. To address these, we introduce the Spiking Structured State Space Model (Spiking-S4). This approach merges the energy efficiency of Spiking Neural Networks (SNN) with the long-range sequence modeling capabilities of Structured State Space Models (S4), offering a compelling solution. Evaluation on the DNS Challenge and VoiceBank+Demand Datasets confirms that Spiking-S4 rivals existing Artificial Neural Network (ANN) methods but with fewer computational resources, as evidenced by reduced parameters and Floating Point Operations (FLOPs).

Keywords

Cite

@article{arxiv.2309.03641,
  title  = {Spiking Structured State Space Model for Monaural Speech Enhancement},
  author = {Yu Du and Xu Liu and Yansong Chua},
  journal= {arXiv preprint arXiv:2309.03641},
  year   = {2024}
}
R2 v1 2026-06-28T12:15:12.530Z