English

Burst2Vec: An Adversarial Multi-Task Approach for Predicting Emotion, Age, and Origin from Vocal Bursts

Sound 2022-10-19 v2 Computation and Language Audio and Speech Processing

Abstract

We present Burst2Vec, our multi-task learning approach to predict emotion, age, and origin (i.e., native country/language) from vocal bursts. Burst2Vec utilises pre-trained speech representations to capture acoustic information from raw waveforms and incorporates the concept of model debiasing via adversarial training. Our models achieve a relative 30 % performance gain over baselines using pre-extracted features and score the highest amongst all participants in the ICML ExVo 2022 Multi-Task Challenge.

Keywords

Cite

@article{arxiv.2206.12469,
  title  = {Burst2Vec: An Adversarial Multi-Task Approach for Predicting Emotion, Age, and Origin from Vocal Bursts},
  author = {Atijit Anuchitanukul and Lucia Specia},
  journal= {arXiv preprint arXiv:2206.12469},
  year   = {2022}
}
R2 v1 2026-06-24T12:03:29.994Z