利用多模态场景信息进行上下文丰富的人体情感感知
计算机视觉与模式识别
2023-03-14 v1 人工智能
计算与语言
摘要
人体情感理解的过程涉及从包括图像、语音和语言在内的各种来源推断特定个体的情绪状态的能力。基于图像的情感感知主要集中于从显著人脸裁剪中提取的表情。然而,人类感知的情绪依赖于多种上下文线索,包括社交环境、前景交互和环境视觉场景。在本工作中,我们利用预训练的视觉-语言(VLN)模型从图像中提取前景上下文描述。此外,我们提出一个多模态上下文融合(MCF)模块,将前景线索与视觉场景及基于人物的上下文信息结合以进行情绪预测。我们在两个与自然场景和电视节目相关的数据集上展示了所提模块化设计的有效性。
引用
@article{arxiv.2303.06904,
title = {Contextually-rich human affect perception using multimodal scene information},
author = {Digbalay Bose and Rajat Hebbar and Krishna Somandepalli and Shrikanth Narayanan},
journal= {arXiv preprint arXiv:2303.06904},
year = {2023}
}
备注
Accepted to IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023