An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
Sound
2024-03-08 v2 Audio and Speech Processing
Abstract
This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts different extraction strategies based on the audio quality, striking a balance between interference removal and speech preservation, which benifits the back-end automatic speech recognition (ASR) systems. Experiments show that our approach achieves a character error rate (CER) of 24.2% and 33.2% on the Dev and Eval set, respectively, obtaining the second place in the challenge.
Keywords
Cite
@article{arxiv.2401.03697,
title = {An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge},
author = {Runduo Han and Xiaopeng Yan and Weiming Xu and Pengcheng Guo and Jiayao Sun and He Wang and Quan Lu and Ning Jiang and Lei Xie},
journal= {arXiv preprint arXiv:2401.03697},
year = {2024}
}
Comments
Accepted by ICASSP 2024