English

VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording

Audio and Speech Processing 2021-07-16 v1 Computation and Language Sound

Abstract

In this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD), based on monotonic chunkwise attention (MoChA) with an auxiliary connectionist temporal classification (CTC) objective. We propose a block-synchronous beam search decoding to take advantage of efficient batched output-synchronous and low-latency input-synchronous searches. We also propose a VAD-free inference algorithm that leverages CTC probabilities to determine a suitable timing to reset the model states to tackle the vulnerability to long-form data. Experimental evaluations demonstrate that the block-synchronous decoding achieves comparable accuracy to the label-synchronous one. Moreover, the VAD-free inference can recognize long-form speech robustly for up to a few hours.

Keywords

Cite

@article{arxiv.2107.07509,
  title  = {VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording},
  author = {Hirofumi Inaguma and Tatsuya Kawahara},
  journal= {arXiv preprint arXiv:2107.07509},
  year   = {2021}
}

Comments

Accepted at Interspeech 2021