English

Human Transcription Quality Improvement

Computation and Language 2023-09-27 v1 Artificial Intelligence Machine Learning Sound Audio and Speech Processing

Abstract

High quality transcription data is crucial for training automatic speech recognition (ASR) systems. However, the existing industry-level data collection pipelines are expensive to researchers, while the quality of crowdsourced transcription is low. In this paper, we propose a reliable method to collect speech transcriptions. We introduce two mechanisms to improve transcription quality: confidence estimation based reprocessing at labeling stage, and automatic word error correction at post-labeling stage. We collect and release LibriCrowd - a large-scale crowdsourced dataset of audio transcriptions on 100 hours of English speech. Experiment shows the Transcription WER is reduced by over 50%. We further investigate the impact of transcription error on ASR model performance and found a strong correlation. The transcription quality improvement provides over 10% relative WER reduction for ASR models. We release the dataset and code to benefit the research community.

Keywords

Cite

@article{arxiv.2309.14372,
  title  = {Human Transcription Quality Improvement},
  author = {Jian Gao and Hanbo Sun and Cheng Cao and Zheng Du},
  journal= {arXiv preprint arXiv:2309.14372},
  year   = {2023}
}

Comments

5 pages, 3 figures, 5 tables, INTERSPEECH 2023

R2 v1 2026-06-28T12:31:56.751Z