English

Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

Computation and Language 2025-03-25 v2 Artificial Intelligence Machine Learning Audio and Speech Processing

Abstract

Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.

Keywords

Cite

@article{arxiv.2502.07029,
  title  = {Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment},
  author = {Kwanghee Choi and Eunjung Yeo and Kalvin Chang and Shinji Watanabe and David Mortensen},
  journal= {arXiv preprint arXiv:2502.07029},
  year   = {2025}
}

Comments

Accepted to NAACL 2025. Codebase available at https://github.com/juice500ml/acoustic-units-for-ood