English

Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding

Computation and Language 2023-05-24 v1 Sound Audio and Speech Processing

Abstract

Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and semantic features in decoder, which is crucial for generating more accurate and informative semantic states. In this paper, we propose an Acoustic and Semantic Cooperative Decoder (ASCD) for ASR. In particular, unlike vanilla decoders that process acoustic and semantic features in two separate stages, ASCD integrates them cooperatively. To prevent information leakage during training, we design a Causal Multimodal Mask. Moreover, a variant Semi-ASCD is proposed to balance accuracy and computational cost. Our proposal is evaluated on the publicly available AISHELL-1 and aidatatang_200zh datasets using Transformer, Conformer, and Branchformer as encoders, respectively. The experimental results show that ASCD significantly improves the performance by leveraging both the acoustic and semantic information cooperatively.

Keywords

Cite

@article{arxiv.2305.14049,
  title  = {Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding},
  author = {Tian-Hao Zhang and Hai-Bo Qin and Zhi-Hao Lai and Song-Lu Chen and Qi Liu and Feng Chen and Xinyuan Qian and Xu-Cheng Yin},
  journal= {arXiv preprint arXiv:2305.14049},
  year   = {2023}
}

Comments

Accepted by Interspeech 2023

R2 v1 2026-06-28T10:42:59.118Z