利用可解释多模态动态注意力融合网络通过远程采集视频预测心境障碍症状
机器学习
2021-09-08 v1
摘要
我们开发了一种新颖的、可解释的多模态分类方法,利用智能手机应用采集的音频、视频和文本来识别心境障碍症状,即抑郁、焦虑与快感缺失。我们使用基于CNN的单模态编码器为每种模态学习动态嵌入,随后通过transformer编码器将这些嵌入进行融合。我们将这些方法应用于一个由智能手机应用采集的新颖数据集,该数据集包含3002名参与者最多三个记录时段的数据。与采用静态嵌入的现有方法相比,我们的方法展现出更优的多模态分类性能。最后,我们使用SHapley Additive exPlanations (SHAP)来优先排序模型中可作为潜在数字标记的重要特征。
引用
@article{arxiv.2109.03029,
title = {Predicting Mood Disorder Symptoms with Remotely Collected Videos Using an Interpretable Multimodal Dynamic Attention Fusion Network},
author = {Tathagata Banerjee and Matthew Kollada and Pablo Gersberg and Oscar Rodriguez and Jane Tiller and Andrew E Jaffe and John Reynders},
journal= {arXiv preprint arXiv:2109.03029},
year = {2021}
}
备注
8 pages, 3 figures, Published in the Computational Approaches to Mental Health Workshop of the International Conference on Machine Learning 2021, https://sites.google.com/view/ca2mh/accepted-papers