English

Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning

Computation and Language 2024-08-16 v3 Artificial Intelligence

Abstract

The majority of Chinese characters are monophonic, while a special group of characters, called polyphonic characters, have multiple pronunciations. As a prerequisite of performing speech-related generative tasks, the correct pronunciation must be identified among several candidates. This process is called Polyphone Disambiguation. Although the problem has been well explored with both knowledge-based and learning-based approaches, it remains challenging due to the lack of publicly available labeled datasets and the irregular nature of polyphone in Mandarin Chinese. In this paper, we propose a novel semi-supervised learning (SSL) framework for Mandarin Chinese polyphone disambiguation that can potentially leverage unlimited unlabeled text data. We explore the effect of various proxy labeling strategies including entropy-thresholding and lexicon-based labeling. Qualitative and quantitative experiments demonstrate that our method achieves state-of-the-art performance. In addition, we publish a novel dataset specifically for the polyphone disambiguation task to promote further research.

Keywords

Cite

@article{arxiv.2102.00621,
  title  = {Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning},
  author = {Yi Shi and Congyi Wang and Yu Chen and Bin Wang},
  journal= {arXiv preprint arXiv:2102.00621},
  year   = {2024}
}
R2 v1 2026-06-23T22:42:34.421Z