English

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

Audio and Speech Processing 2026-03-03 v1

Abstract

This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike conventional meetings centered on a single shared topic, this scenario contains multiple parallel dialogues--up to eight speakers across up to four simultaneous conversations--with a speech overlap rate exceeding 90%. To tackle this, we propose a multimodal cascaded system that leverages per-speaker visual streams extracted from synchronized 360 degree video together with single-channel audio. Our system improves three components of the pipeline by leveraging enhanced audio-visual pretrained models: Active Speaker Detection (ASD), Audio-Visual Target Speech Extraction (AVTSE), and Audio-Visual Speech Recognition (AVSR). The AVSR module further incorporates Whisper and LLM techniques to boost transcription accuracy. Our best single cascaded system achieves a Speaker Word Error Rate (WER) of 32.44% on the development set. By further applying ROVER to fuse outputs from diverse front-end and back-end variants, we reduce Speaker WER to 31.40%. Notably, our LLM-based zero-shot conversational clustering achieves a speaker clustering F1 score of 1.0, yielding a final Joint ASR-Clustering Error Rate (JACER) of 15.70%.

Keywords

Cite

@article{arxiv.2603.01415,
  title  = {The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge},
  author = {Ya Jiang and Ruoyu Wang and Jingxuan Zhang and Jun Du and Yi Han and Zihao Quan and Hang Chen and Yeran Yang and Kongzhi Zheng and Zhuo Chen and Yanhui Tu and Shutong Niu and Changfeng Xi and Mengzhi Wang and Zhongbin Wu and Jieru Chen and Henghui Zhi and Weiyi Shi and Shuhang Wu and Genshun Wan and Jia Pan and Jianqing Gao},
  journal= {arXiv preprint arXiv:2603.01415},
  year   = {2026}
}
R2 v1 2026-07-01T10:58:28.230Z