普通话 ASR 中识别与转写的解耦
计算与语言
2021-08-04 v1 声音
音频与语音处理
摘要
近期关于自动语音识别(ASR)的文献大多采用端到端方法。与书写系统与发音密切相关的英语不同,汉字(Hanzi)表意而非表音。我们提出将音频→汉字分解为两个子任务:(1)音频→拼音和(2)拼音→汉字,其中拼音是标准汉语的音标系统。以这种方式分解音频→汉字任务,在 Aishell-1 语料库上实现了 3.9% 的字符错误率(CER),这是迄今为止在该数据集上报告的最佳结果。
引用
@article{arxiv.2108.01129,
title = {Decoupling recognition and transcription in Mandarin ASR},
author = {Jiahong Yuan and Xingyu Cai and Dongji Gao and Renjie Zheng and Liang Huang and Kenneth Church},
journal= {arXiv preprint arXiv:2108.01129},
year = {2021}
}
备注
submitted to ASRU 2021