中文

基于通用音素识别和语言特定音素对比建模的多语言嗓性失调语音评估

计算与语言 2026-02-12 v2 声音 音频与语音处理

摘要

The growing prevalence of neurological disorders associated with dysarthria motivates the need for automated intelligibility assessment methods that are applicalbe across languages. However, most existing approaches are either limited to a single language or fail to capture language-specific factors shaping intelligibility. We present a multilingual phoneme-production assessment framework that integrates universal phone recognition with language-specific phoneme interpretation using contrastive phonological feature distances for phone-to-phoneme mapping and sequence alignment. The framework yields three metrics: phoneme error rate (PER), phonological feature error rate (PFER), and a newly proposed alignment-free measure, phoneme coverage (PhonCov). Analysis on English, Spanish, Italian, and Tamil show that PER benefits from the combination of mapping and alignment, PFER from alignment alone, and PhonCov from mapping. Further analyses demonstrate that the proposed framework captures clinically meaningful patterns of intelligibility degradation consistent with established observations of dysarthric speech.

关键词

引用

@article{arxiv.2601.21205,
  title  = {Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling},
  author = {Eunjung Yeo and Julie M. Liss and Visar Berisha and David R. Mortensen},
  journal= {arXiv preprint arXiv:2601.21205},
  year   = {2026}
}

备注

10 pages, 4 figures