English

PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting

Audio and Speech Processing 2026-03-20 v1 Artificial Intelligence Computation and Language Sound

Abstract

As advancements in technologies like Internet of Things (IoT), Automatic Speech Recognition (ASR), Speaker Verification (SV), and Text-to-Speech (TTS) lead to increased usage of intelligent voice assistants, the demand for privacy and personalization has escalated. In this paper, we introduce a multi-task learning framework for personalized, customizable open-vocabulary Keyword Spotting (PCOV-KWS). This framework employs a lightweight network to simultaneously perform Keyword Spotting (KWS) and SV to address personalized KWS requirements. We have integrated a training criterion distinct from softmax-based loss, transforming multi-class classification into multiple binary classifications, which eliminates inter-category competition, while an optimization strategy for multi-task loss weighting is employed during training. We evaluated our PCOV-KWS system in multiple datasets, demonstrating that it outperforms the baselines in evaluation results, while also requiring fewer parameters and lower computational resources.

Keywords

Cite

@article{arxiv.2603.18023,
  title  = {PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting},
  author = {Jianan Pan and Kejie Huang},
  journal= {arXiv preprint arXiv:2603.18023},
  year   = {2026}
}
R2 v1 2026-07-01T11:26:44.408Z