English

ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody

Audio and Speech Processing 2026-03-20 v1 Artificial Intelligence Computation and Language Sound

Abstract

Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.

Keywords

Cite

@article{arxiv.2603.18024,
  title  = {ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody},
  author = {Jianan Pan and Yuanming Zhang and Kejie Huang},
  journal= {arXiv preprint arXiv:2603.18024},
  year   = {2026}
}