English

Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition

Sound 2026-03-30 v2

Abstract

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data.

Keywords

Cite

@article{arxiv.2603.25377,
  title  = {Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition},
  author = {Yuhang Dai and Haopeng Lin and Jiale Qian and Ruiqi Yan and Hao Meng and Hanke Xie and Hanlin Wen and Shunshun Yin and Ming Tao and Xie Chen and Lei Xie and Xinsheng Wang},
  journal= {arXiv preprint arXiv:2603.25377},
  year   = {2026}
}

Comments

5 pages, 2 figures, 2 tables

R2 v1 2026-07-01T11:39:09.819Z