English

CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition

Computation and Language 2020-02-13 v4 Machine Learning Neural and Evolutionary Computing Sound Audio and Speech Processing

Abstract

In this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction. It is inspired by the integrate-and-fire model in spiking neural networks and employed in the encoder-decoder framework consists of continuous functions, thus being named as: Continuous Integrate-and-Fire (CIF). Applied to the ASR task, CIF not only shows a concise calculation, but also supports online recognition and acoustic boundary positioning, thus suitable for various ASR scenarios. Several support strategies are also proposed to alleviate the unique problems of CIF-based model. With the joint action of these methods, the CIF-based model shows competitive performance. Notably, it achieves a word error rate (WER) of 2.86% on the test-clean of Librispeech and creates new state-of-the-art result on Mandarin telephone ASR benchmark.

Keywords

Cite

@article{arxiv.1905.11235,
  title  = {CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition},
  author = {Linhao Dong and Bo Xu},
  journal= {arXiv preprint arXiv:1905.11235},
  year   = {2020}
}

Comments

To appear at ICASSP 2020

R2 v1 2026-06-23T09:26:38.862Z