English

Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge

Computer Vision and Pattern Recognition 2026-07-15 v1

Abstract

Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing, and assume each speaker only speaks a single language. However, in real-world applications, such assumptions often do not hold. Visual or audio information may be missing due to occlusions, camera or microphone failures, or privacy constraints. Multilingual speakers introduce additional complexity due to linguistic variability across languages. These situations constitute substantial challenges for the robustness and generalization capabilities of multimodal speaker identification systems. Aim of the POLY-SIM 2026 challenge is to address these aspects of speaker identification and to provide a standardized setup for the comparison of the proposed solutions.

Keywords

Cite

@article{arxiv.2607.13669,
  title  = {Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge},
  author = {Marta Moscati and Muhammad Saad Saeed and Marina Zanoni and Mubashir Noman and Rohan Kumar Das and Monorama Swain and Yassin Terraf and Yufang Hou and Elisabeth Andre and Khalid Mahmood Malik and Markus Schedl and Shah Nawaz},
  journal= {arXiv preprint arXiv:2607.13669},
  year   = {2026}
}

Comments

Accepted at ACM MM 2026