English

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Computation and Language 2026-07-23 v1 Artificial Intelligence

Abstract

Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.

Keywords

Cite

@article{arxiv.2607.21332,
  title  = {Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin},
  author = {Zhiheng Qian and Aini Li and Hai Hu and Liang Zhao},
  journal= {arXiv preprint arXiv:2607.21332},
  year   = {2026}
}

Comments

5 pages, 1 figure