English

Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds

Audio and Speech Processing 2025-09-23 v1 Sound

Abstract

This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse languages, accents, and speech patterns, are particularly adept at navigating the noisy and complex acoustic environments characteristic of crowd settings, thereby offering a significant advantage for CER. To substantiate this, we perform a comprehensive analysis, comparing polyglot, monolingual, and speaker recognition SFMs through extensive experiments on a benchmark CER dataset across varying audio durations (1 sec, 500 ms, and 250 ms). The results consistently demonstrate the superiority of polyglot SFMs, outperforming their counterparts across all audio lengths and excelling even with extremely short-duration inputs. These findings pave the way for adaptation of SFMs in setting up new benchmarks for CER.

Keywords

Cite

@article{arxiv.2509.16329,
  title  = {Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds},
  author = {Orchid Chetia Phukan and Girish and Mohd Mujtaba Akhtar and Panchal Nayak and Priyabrata Mallick and Swarup Ranjan Behera and Parabattina Bhagath and Pailla Balakrishna Reddy and Arun Balaji Buduru},
  journal= {arXiv preprint arXiv:2509.16329},
  year   = {2025}
}

Comments

Accepted to APSIPA-ASC 2025

R2 v1 2026-07-01T05:46:31.890Z