English

AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Computation and Language 2018-09-14 v2

Abstract

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research. It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR. In AISHELL-2, 1000 hours of clean read-speech data from iOS is published, which is free for academic usage. On top of AISHELL-2 corpus, an improved recipe is developed and released, containing key components for industrial applications, such as Chinese word segmentation, flexible vocabulary expension and phone set transformation etc. Pipelines support various state-of-the-art techniques, such as time-delayed neural networks and Lattic-Free MMI objective funciton. In addition, we also release dev and test data from other channels(Android and Mic). For research community, we hope that AISHELL-2 corpus can be a solid resource for topics like transfer learning and robust ASR. For industry, we hope AISHELL-2 recipe can be a helpful reference for building meaningful industrial systems and products.

Cite

@article{arxiv.1808.10583,
  title  = {AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale},
  author = {Jiayu Du and Xingyu Na and Xuechen Liu and Hui Bu},
  journal= {arXiv preprint arXiv:1808.10583},
  year   = {2018}
}
R2 v1 2026-06-23T03:49:59.021Z