English

Low-activity supervised convolutional spiking neural networks applied to speech commands recognition

Machine Learning 2020-11-16 v1 Computation and Language Sound Audio and Speech Processing

Abstract

Deep Neural Networks (DNNs) are the current state-of-the-art models in many speech related tasks. There is a growing interest, though, for more biologically realistic, hardware friendly and energy efficient models, named Spiking Neural Networks (SNNs). Recently, it has been shown that SNNs can be trained efficiently, in a supervised manner, using backpropagation with a surrogate gradient trick. In this work, we report speech command (SC) recognition experiments using supervised SNNs. We explored the Leaky-Integrate-Fire (LIF) neuron model for this task, and show that a model comprised of stacked dilated convolution spiking layers can reach an error rate very close to standard DNNs on the Google SC v1 dataset: 5.5%, while keeping a very sparse spiking activity, below 5%, thank to a new regularization term. We also show that modeling the leakage of the neuron membrane potential is useful, since the LIF model outperformed its non-leaky model counterpart significantly.

Keywords

Cite

@article{arxiv.2011.06846,
  title  = {Low-activity supervised convolutional spiking neural networks applied to speech commands recognition},
  author = {Thomas Pellegrini and Romain Zimmer and Timothée Masquelier},
  journal= {arXiv preprint arXiv:2011.06846},
  year   = {2020}
}

Comments

Accepted to IEEE Spoken Language Technology Workshop 2021

R2 v1 2026-06-23T20:10:23.988Z