English

Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition

Sound 2024-07-04 v1 Artificial Intelligence Audio and Speech Processing

Abstract

Currently, end-to-end (E2E) speech recognition methods have achieved promising performance. However, auto speech recognition (ASR) models still face challenges in recognizing multi-accent speech accurately. We propose a layer-adapted fusion (LAF) model, called Qifusion-Net, which does not require any prior knowledge about the target accent. Based on dynamic chunk strategy, our approach enables streaming decoding and can extract frame-level acoustic feature, facilitating fine-grained information fusion. Experiment results demonstrate that our proposed methods outperform the baseline with relative reductions of 22.1%\% and 17.2%\% in character error rate (CER) across multi accent test datasets on KeSpeech and MagicData-RMAC.

Keywords

Cite

@article{arxiv.2407.03026,
  title  = {Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition},
  author = {Jinming Chen and Jingyi Fang and Yuanzhong Zheng and Yaoxuan Wang and Haojun Fei},
  journal= {arXiv preprint arXiv:2407.03026},
  year   = {2024}
}

Comments

accpeted by interspeech 2014, 5 pages, 1 figure

R2 v1 2026-06-28T17:27:48.706Z