中文

AICA-Bench:全面检视视觉语言模型在情感图像内容分析中的能力

计算机视觉与模式识别 2026-04-08 v1

摘要

视觉语言模型(Vision-Language Models, VLMs)在感知任务中展现出强大的能力,但综合感知、推理与生成于统一框架中的情感图像内容分析(Affective Image Content Analysis, AICA)仍受到 insufficient 探索。为填补这一空白,我们引入 AICA-Bench——一个包含三项核心任务的全面基准:情感理解(Emotion Understanding, EU)、情感推理(Emotion Reasoning, ER)和情感引导的内容生成(Emotion-Guided Content Generation, EGCG)。我们评估了 23 个视觉语言模型,识别出两个主要局限:情感强度校准不足和浅层开放式描述。为此,我们提出了 Grounded Affective Tree(GAT)提示框架,该框架结合视觉 scaffolding 与层次化推理。实验表明,GAT 减少了情感强度误差,提升了描述深度,为未来在情感多模态理解与生成方面的研究提供了强有力的基线。

关键词

引用

@article{arxiv.2604.05900,
  title  = {AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis},
  author = {Dong She and Xianrong Yao and Liqun Chen and Jinghe Yu and Yang Gao and Zhanpeng Jin},
  journal= {arXiv preprint arXiv:2604.05900},
  year   = {2026}
}

备注

Accepted by Findings of ACL 2026