English

The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022

Sound 2023-07-31 v1 Audio and Speech Processing

Abstract

This paper describes the FlySpeech speaker diarization system submitted to the second \textbf{M}ultimodal \textbf{I}nformation Based \textbf{S}peech \textbf{P}rocessing~(\textbf{MISP}) Challenge held in ICASSP 2022. We develop an end-to-end audio-visual speaker diarization~(AVSD) system, which consists of a lip encoder, a speaker encoder, and an audio-visual decoder. Specifically, to mitigate the degradation of diarization performance caused by separate training, we jointly train the speaker encoder and the audio-visual decoder. In addition, we leverage the large-data pretrained speaker extractor to initialize the speaker encoder.

Keywords

Cite

@article{arxiv.2307.15400,
  title  = {The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022},
  author = {Li Zhang and Huan Zhao and Yue Li and Bowen Pang and Yannan Wang and Hongji Wang and Wei Rao and Qing Wang and Lei Xie},
  journal= {arXiv preprint arXiv:2307.15400},
  year   = {2023}
}
R2 v1 2026-06-28T11:42:40.421Z