English

Sequence-to-sequence Models for Small-Footprint Keyword Spotting

Sound 2018-11-02 v1 Audio and Speech Processing

Abstract

In this paper, we propose a sequence-to-sequence model for keyword spotting (KWS). Compared with other end-to-end architectures for KWS, our model simplifies the pipelines of production-quality KWS system and satisfies the requirement of high accuracy, low-latency, and small-footprint. We also evaluate the performances of different encoder architectures, which include LSTM and GRU. Experiments on the real-world wake-up data show that our approach outperforms the recently proposed attention-based end-to-end model. Specifically speaking, with 73K parameters, our sequence-to-sequence model achieves \sim3.05\% false rejection rate (FRR) at 0.1 false alarm (FA) per hour.

Keywords

Cite

@article{arxiv.1811.00348,
  title  = {Sequence-to-sequence Models for Small-Footprint Keyword Spotting},
  author = {Haitong Zhang and Junbo Zhang and Yujun Wang},
  journal= {arXiv preprint arXiv:1811.00348},
  year   = {2018}
}

Comments

Submitted to ICASSP 2019

R2 v1 2026-06-23T05:00:31.075Z