English

BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR

Sound 2023-10-06 v3 Computation and Language Audio and Speech Processing

Abstract

The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can make speaker change prediction difficult. To address this, we propose boundary-aware serialized output training (BA-SOT), which explicitly incorporates boundary knowledge into the decoder via a speaker change detection task and boundary constraint loss. We also introduce a two-stage connectionist temporal classification (CTC) strategy that incorporates token-level SOT CTC to restore temporal context information. Besides typical character error rate (CER), we introduce utterance-dependent character error rate (UD-CER) to further measure the precision of speaker change prediction. Compared to original SOT, BA-SOT reduces CER/UD-CER by 5.1%/14.0%, and leveraging a pre-trained ASR model for BA-SOT model initialization further reduces CER/UD-CER by 8.4%/19.9%.

Keywords

Cite

@article{arxiv.2305.13716,
  title  = {BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR},
  author = {Yuhao Liang and Fan Yu and Yangze Li and Pengcheng Guo and Shiliang Zhang and Qian Chen and Lei Xie},
  journal= {arXiv preprint arXiv:2305.13716},
  year   = {2023}
}

Comments

Accepted by INTERSPEECH 2023

R2 v1 2026-06-28T10:42:28.915Z