Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
Machine Learning
2025-04-28 v2 Computation and Language
Sound
Audio and Speech Processing
Abstract
We propose using cascaded classifiers for a keyword spotting (KWS) task on narrow-band (NB), 8kHz audio acquired in non-IID environments -- a more challenging task than most state-of-the-art KWS systems face. We present a model that incorporates Deep Neural Networks (DNNs), cascading, multiple-feature representations, and multiple-instance learning. The cascaded classifiers handle the task's class imbalance and reduce power consumption on computationally-constrained devices via early termination. The KWS system achieves a false negative rate of 6% at an hourly false positive rate of 0.75
Cite
@article{arxiv.1711.08058,
title = {Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio},
author = {Ahmad AbdulKader and Kareem Nassar and Mohamed El-Geish and Daniel Galvez and Chetan Patil},
journal= {arXiv preprint arXiv:1711.08058},
year = {2025}
}
Comments
Published in the proceedings of NeurIPS 2017 Workshop: Machine Learning on the Phone and other Consumer Devices