English

Thai Wav2Vec2.0 with CommonVoice V8

Computation and Language 2022-08-10 v1 Sound Audio and Speech Processing

Abstract

Recently, Automatic Speech Recognition (ASR), a system that converts audio into text, has caught a lot of attention in the machine learning community. Thus, a lot of publicly available models were released in HuggingFace. However, most of these ASR models are available in English; only a minority of the models are available in Thai. Additionally, most of the Thai ASR models are closed-sourced, and the performance of existing open-sourced models lacks robustness. To address this problem, we train a new ASR model on a pre-trained XLSR-Wav2Vec model with the Thai CommonVoice corpus V8 and train a trigram language model to boost the performance of our ASR model. We hope that our models will be beneficial to individuals and the ASR community in Thailand.

Keywords

Cite

@article{arxiv.2208.04799,
  title  = {Thai Wav2Vec2.0 with CommonVoice V8},
  author = {Wannaphong Phatthiyaphaibun and Chompakorn Chaksangchaichot and Peerat Limkonchotiwat and Ekapol Chuangsuwanich and Sarana Nutanong},
  journal= {arXiv preprint arXiv:2208.04799},
  year   = {2022}
}
R2 v1 2026-06-25T01:35:56.869Z