中文

基于可微K-Means的L1-L2多任务学习中的高级Interlanguage Speech Intelligibility Benefit建模:面向语音识别中带重音的离散标记

声音 2026-01-28 v1

摘要

构建对外语重音语音鲁棒的ASR系统是当今全球化时代的重要挑战。先前的研究探索了通过使用说话人L1学习k-means聚类质心在SSL特征空间中获取语音标记,以实现对重音语音的语音识别性能提升的方法,复现了已知的Interlanguage Speech Intelligibility Benefit(ISIB)现象,即带外语重音的语音对共享说话人母语的听者比对母语听者更易被理解。ISIB通过使用说话人的L1学习k-means聚类中心来实现。本研究提出了更先进的ISIB建模方法。通过采用可微k-means并针对L1和L2 ASR优化整个模块,所提出的方法在仅使用本土语音以及在有限 amount of accented speech(带外语重音语音)两种情况下均优于基线方法。值得注意的是,在后者场景下,我们的方法实现了约20%的识别准确率提升。

关键词

引用

@article{arxiv.2601.19767,
  title  = {Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR},
  author = {Kentaro Onda and Satoru Fukayama and Daisuke Saito and Nobuaki Minematsu},
  journal= {arXiv preprint arXiv:2601.19767},
  year   = {2026}
}

备注

Accepted to ICASSP 2026