English

Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection

Audio and Speech Processing 2024-09-24 v1 Sound

Abstract

In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection.

Keywords

Cite

@article{arxiv.2409.14312,
  title  = {Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection},
  author = {Orchid Chetia Phukan and Swarup Ranjan Behera and Shubham Singh and Muskaan Singh and Vandana Rajan and Arun Balaji Buduru and Rajesh Sharma and S. R. Mahadeva Prasanna},
  journal= {arXiv preprint arXiv:2409.14312},
  year   = {2024}
}

Comments

Submitted to ICASSP 2025