English

iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis

Audio and Speech Processing 2026-02-26 v1 Computation and Language

Abstract

This work presents iMiGUE-Speech, an extension of the iMiGUE dataset that provides a spontaneous affective corpus for studying emotional and affective states. The new release focuses on speech and enriches the original dataset with additional metadata, including speech transcripts, speaker-role separation between interviewer and interviewee, and word-level forced alignments. Unlike existing emotional speech datasets that rely on acted or laboratory-elicited emotions, iMiGUE-Speech captures spontaneous affect arising naturally from real match outcomes. To demonstrate the utility of the dataset and establish initial benchmarks, we introduce two evaluation tasks for comparative assessment: speech emotion recognition and transcript-based sentiment analysis. These tasks leverage state-of-the-art pre-trained representations to assess the dataset's ability to capture spontaneous affective states from both acoustic and linguistic modalities. iMiGUE-Speech can also be synchronously paired with micro-gesture annotations from the original iMiGUE dataset, forming a uniquely multimodal resource for studying speech-gesture affective dynamics. The extended dataset is available at https://github.com/CV-AC/imigue-speech.

Keywords

Cite

@article{arxiv.2602.21464,
  title  = {iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis},
  author = {Sofoklis Kakouros and Fang Kang and Haoyu Chen},
  journal= {arXiv preprint arXiv:2602.21464},
  year   = {2026}
}

Comments

Accepted to Speech Prosody 2026

R2 v1 2026-07-01T10:50:54.724Z