English

Decoupling recognition and transcription in Mandarin ASR

Computation and Language 2021-08-04 v1 Sound Audio and Speech Processing

Abstract

Much of the recent literature on automatic speech recognition (ASR) is taking an end-to-end approach. Unlike English where the writing system is closely related to sound, Chinese characters (Hanzi) represent meaning, not sound. We propose factoring audio -> Hanzi into two sub-tasks: (1) audio -> Pinyin and (2) Pinyin -> Hanzi, where Pinyin is a system of phonetic transcription of standard Chinese. Factoring the audio -> Hanzi task in this way achieves 3.9% CER (character error rate) on the Aishell-1 corpus, the best result reported on this dataset so far.

Keywords

Cite

@article{arxiv.2108.01129,
  title  = {Decoupling recognition and transcription in Mandarin ASR},
  author = {Jiahong Yuan and Xingyu Cai and Dongji Gao and Renjie Zheng and Liang Huang and Kenneth Church},
  journal= {arXiv preprint arXiv:2108.01129},
  year   = {2021}
}

Comments

submitted to ASRU 2021

R2 v1 2026-06-24T04:46:09.957Z