English

Quaternion Neural Networks for Multi-channel Distant Speech Recognition

Audio and Speech Processing 2020-05-20 v2 Machine Learning Sound Machine Learning

Abstract

Despite the significant progress in automatic speech recognition (ASR), distant ASR remains challenging due to noise and reverberation. A common approach to mitigate this issue consists of equipping the recording devices with multiple microphones that capture the acoustic scene from different perspectives. These multi-channel audio recordings contain specific internal relations between each signal. In this paper, we propose to capture these inter- and intra- structural dependencies with quaternion neural networks, which can jointly process multiple signals as whole quaternion entities. The quaternion algebra replaces the standard dot product with the Hamilton one, thus offering a simple and elegant way to model dependencies between elements. The quaternion layers are then coupled with a recurrent neural network, which can learn long-term dependencies in the time domain. We show that a quaternion long-short term memory neural network (QLSTM), trained on the concatenated multi-channel speech signals, outperforms equivalent real-valued LSTM on two different tasks of multi-channel distant speech recognition.

Keywords

Cite

@article{arxiv.2005.08566,
  title  = {Quaternion Neural Networks for Multi-channel Distant Speech Recognition},
  author = {Xinchi Qiu and Titouan Parcollet and Mirco Ravanelli and Nicholas Lane and Mohamed Morchid},
  journal= {arXiv preprint arXiv:2005.08566},
  year   = {2020}
}

Comments

4 pages

R2 v1 2026-06-23T15:37:10.730Z