English

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition

Computation and Language 2025-01-30 v1 Sound Audio and Speech Processing

Abstract

We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) that enhances multilingual performance, especially for low-resource languages. It utilizes a cross-lingual embedding clustering method to construct a hierarchical Softmax (H-Softmax) decoder, which enables similar tokens across different languages to share similar decoder representations. It addresses the limitations of the previous Huffman-based H-Softmax method, which relied on shallow features in token similarity assessments. Through experiments on a downsampled dataset of 15 languages, we demonstrate the effectiveness of our approach in improving low-resource multilingual ASR accuracy.

Keywords

Cite

@article{arxiv.2501.17615,
  title  = {Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition},
  author = {Zhengdong Yang and Qianying Liu and Sheng Li and Fei Cheng and Chenhui Chu},
  journal= {arXiv preprint arXiv:2501.17615},
  year   = {2025}
}
R2 v1 2026-06-28T21:23:43.613Z