中文

为印度语言构建稳健且可扩展的多语言语音识别系统

计算与语言 2025-11-20 v1 人工智能

摘要

本文描述了SPRING实验室、印度理工学院玛塔纳斯拉姆大学为ASRU MADASR 2.0挑战所开发的系统。所开发的系统侧重于改进语音识别系统对8种语言及其33个方言的预测能力。我们参与了第1轨和第2轨,这两个轨道限制使用额外数据并从零开发多语言系统。我们提出了一种新颖的训练方法,使用带有音素公共标签集(CLS)作为中间表示的Multi-Decoder架构。它在CLS空间中的表现优于基线。我们还讨论了 various methods used to retain the gain obtained in the phonemic space while converting them back to the corresponding grapheme representations. Our systems beat the baseline in 3 languages (Track 2) in terms of WER/CER and achieved the highest language ID and dialect ID accuracy among all participating teams (Track 2).

关键词

引用

@article{arxiv.2511.15418,
  title  = {Building Robust and Scalable Multilingual ASR for Indian Languages},
  author = {Arjun Gangwar and Kaousheik Jayakumar and S. Umesh},
  journal= {arXiv preprint arXiv:2511.15418},
  year   = {2025}
}