English

Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System

Audio and Speech Processing 2024-05-21 v1

Abstract

This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, front-end speech dereverberation, beamforming and robust i-vector extraction with comparisons of our in-house implementations and publicly available tools. We finally achieved a word error rate of 69.4% on the development set, which is a 11.7% absolute improvement over the previous baseline of 81.1%, and release this improved baseline with refined techniques/tools as an advanced CHiME-5 recipe.

Keywords

Cite

@article{arxiv.2405.11078,
  title  = {Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System},
  author = {Vimal Manohar and Szu-Jui Chen and Zhiqi Wang and Yusuke Fujita and Shinji Watanabe and Sanjeev Khudanpur},
  journal= {arXiv preprint arXiv:2405.11078},
  year   = {2024}
}

Comments

Published in: ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

R2 v1 2026-06-28T16:31:28.100Z