English

Efficiently Adapting Spoken Language Models for the Singaporean Context

Computation and Language 2026-07-11 v1 Artificial Intelligence

Abstract

Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HT-Moonstone (5B) matches or outperforms SLMs up to 7x its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2\% of its original speech QA ability.

Cite

@article{arxiv.2607.10092,
  title  = {Efficiently Adapting Spoken Language Models for the Singaporean Context},
  author = {Ng Jia Sheng Jason},
  journal= {arXiv preprint arXiv:2607.10092},
  year   = {2026}
}

Comments

10 pages, 2 figures

R2 v1 2026-07-22T20:35:50.650Z