English

A Fast-Converged Acoustic Modeling for Korean Speech Recognition: A Preliminary Study on Time Delay Neural Network

Computation and Language 2018-07-17 v1 Sound Audio and Speech Processing

Abstract

In this paper, a time delay neural network (TDNN) based acoustic model is proposed to implement a fast-converged acoustic modeling for Korean speech recognition. The TDNN has an advantage in fast-convergence where the amount of training data is limited, due to subsampling which excludes duplicated weights. The TDNN showed an absolute improvement of 2.12% in terms of character error rate compared to feed forward neural network (FFNN) based modelling for Korean speech corpora. The proposed model converged 1.67 times faster than a FFNN-based model did.

Keywords

Cite

@article{arxiv.1807.05855,
  title  = {A Fast-Converged Acoustic Modeling for Korean Speech Recognition: A Preliminary Study on Time Delay Neural Network},
  author = {Hosung Park and Donghyun Lee and Minkyu Lim and Yoseb Kang and Juneseok Oh and Ji-Hwan Kim},
  journal= {arXiv preprint arXiv:1807.05855},
  year   = {2018}
}

Comments

6 pages, 2 figures