English

Self-supervision and Learnable STRFs for Age, Emotion, and Country Prediction

Sound 2022-06-28 v1 Artificial Intelligence Audio and Speech Processing

Abstract

This work presents a multitask approach to the simultaneous estimation of age, country of origin, and emotion given vocal burst audio for the 2022 ICML Expressive Vocalizations Challenge ExVo-MultiTask track. The method of choice utilized a combination of spectro-temporal modulation and self-supervised features, followed by an encoder-decoder network organized in a multitask paradigm. We evaluate the complementarity between the tasks posed by examining independent task-specific and joint models, and explore the relative strengths of different feature sets. We also introduce a simple score fusion mechanism to leverage the complementarity of different feature sets for this task. We find that robust data preprocessing in conjunction with score fusion over spectro-temporal receptive field and HuBERT models achieved our best ExVo-MultiTask test score of 0.412.

Keywords

Cite

@article{arxiv.2206.12568,
  title  = {Self-supervision and Learnable STRFs for Age, Emotion, and Country Prediction},
  author = {Roshan Sharma and Tyler Vuong and Mark Lindsey and Hira Dhamyal and Rita Singh and Bhiksha Raj},
  journal= {arXiv preprint arXiv:2206.12568},
  year   = {2022}
}