English

EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting

Audio and Speech Processing 2026-01-26 v1 Computation and Language Sound

Abstract

We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.

Keywords

Cite

@article{arxiv.2601.16316,
  title  = {EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting},
  author = {Oguzhan Buyuksolak and Alican Gok and Osman Erman Okman},
  journal= {arXiv preprint arXiv:2601.16316},
  year   = {2026}
}

Comments

Accepted to be presented in IEEE ICASSP 2026

R2 v1 2026-07-01T09:16:33.480Z