English

Locale Encoding For Scalable Multilingual Keyword Spotting Models

Computation and Language 2023-02-28 v1 Machine Learning

Abstract

A Multilingual Keyword Spotting (KWS) system detects spokenkeywords over multiple locales. Conventional monolingual KWSapproaches do not scale well to multilingual scenarios because ofhigh development/maintenance costs and lack of resource sharing.To overcome this limit, we propose two locale-conditioned universalmodels with locale feature concatenation and feature-wise linearmodulation (FiLM). We compare these models with two baselinemethods: locale-specific monolingual KWS, and a single universalmodel trained over all data. Experiments over 10 localized languagedatasets show that locale-conditioned models substantially improveaccuracy over baseline methods across all locales in different noiseconditions.FiLMperformed the best, improving on average FRRby 61% (relative) compared to monolingual KWS models of similarsizes.

Keywords

Cite

@article{arxiv.2302.12961,
  title  = {Locale Encoding For Scalable Multilingual Keyword Spotting Models},
  author = {Pai Zhu and Hyun Jin Park and Alex Park and Angelo Scorza Scarpati and Ignacio Lopez Moreno},
  journal= {arXiv preprint arXiv:2302.12961},
  year   = {2023}
}

Comments

Accepted for ICASSP 2023

R2 v1 2026-06-28T08:49:16.834Z