English

Self-Supervised Embeddings for Detecting Individual Symptoms of Depression

Sound 2024-06-26 v1 Machine Learning Audio and Speech Processing

Abstract

Depression, a prevalent mental health disorder impacting millions globally, demands reliable assessment systems. Unlike previous studies that focus solely on either detecting depression or predicting its severity, our work identifies individual symptoms of depression while also predicting its severity using speech input. We leverage self-supervised learning (SSL)-based speech models to better utilize the small-sized datasets that are frequently encountered in this task. Our study demonstrates notable performance improvements by utilizing SSL embeddings compared to conventional speech features. We compare various types of SSL pretrained models to elucidate the type of speech information (semantic, speaker, or prosodic) that contributes the most in identifying different symptoms. Additionally, we evaluate the impact of combining multiple SSL embeddings on performance. Furthermore, we show the significance of multi-task learning for identifying depressive symptoms effectively.

Keywords

Cite

@article{arxiv.2406.17229,
  title  = {Self-Supervised Embeddings for Detecting Individual Symptoms of Depression},
  author = {Sri Harsha Dumpala and Katerina Dikaios and Abraham Nunes and Frank Rudzicz and Rudolf Uher and Sageev Oore},
  journal= {arXiv preprint arXiv:2406.17229},
  year   = {2024}
}

Comments

Accepted at INTERSPEECH 2024

R2 v1 2026-06-28T17:18:11.483Z