English

MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting

Audio and Speech Processing 2024-06-12 v1 Computation and Language Sound

Abstract

In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.

Keywords

Cite

@article{arxiv.2406.07310,
  title  = {MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting},
  author = {Zhiqi Ai and Zhiyong Chen and Shugong Xu},
  journal= {arXiv preprint arXiv:2406.07310},
  year   = {2024}
}

Comments

Accepted at INTERSPEECH 2024

R2 v1 2026-06-28T17:01:37.347Z