TSPC:一种用于越南-英语语种切换语音识别的两阶段音素中心架构
摘要
Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems。 Existing methods often fail to capture the sub tle phonological shifts inherent in CS scenarios。 The challenge is particu larly difficult for language pairs like Vietnamese and English, where both distinct phonological features and the ambiguity arising from similar sound recognition are present。In this paper, we propose a novel architecture for Vietnamese-English CS ASR, a Two-Stage Phoneme-Centric model (TSPC)。 TSPC adopts a phoneme-centric approach based on an extended Vietnamese phoneme set as an intermediate representation for mixed-lingual modeling, while remaining efficient under low computational-resource constraints。Experimental results demonstrate that TSPC consistently outperforms exist ing baselines, including PhoWhisper-base, in Vietnamese-English CS ASR, achieving a significantly lower word error rate of 19.06% with reduced training resources。Furthermore, the phonetic-based two-stage architecture enables phoneme adaptation and language conversion to enhance ASR performance in complex CS Vietnamese-English ASR scenarios。
引用
@article{arxiv.2509.05983,
title = {TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition},
author = {Tran Nguyen Anh and Truong Dinh Dung and Vo Van Nam and Minh N. H. Nguyen},
journal= {arXiv preprint arXiv:2509.05983},
year = {2026}
}
备注
Update new version