English

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

Computer Vision and Pattern Recognition 2026-05-26 v2

Abstract

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still limited to coarse-grained emotion perception or deficient reasoning capabilities. To bridge this gap, we introduce \textbf{EEmoDB}, the largest image-{\ul e}voked {\ul emo}tion understanding {\ul d}ataset to date. It features 55 analysis dimensions spanning 55 distinct task categories, facilitating comprehensive interpretation. Specifically, we compile 1.2M1.2M question-answering (QA) pairs (EEmoDB-QA) from 125K125K images via automated generation, alongside a 36K36K dataset (EEmoDB-Assess) curated from 25K25K images for fine-grained assessment. Furthermore, we propose \textbf{EEmo-Logic}, an \textbf{all-in-one} multimodal large language model (MLLM) developed via instruction fine-tuning and task-customized group relative preference optimization (GRPO) with novel reward design. Extensive experiments demonstrate that EEmo-Logic achieves robust performance in in-domain and cross-domain datasets, excelling in emotion QA and fine-grained assessment. The dataset and code are available at https://github.com/workerred/EEmo-Logic.

Keywords

Cite

@article{arxiv.2602.01173,
  title  = {EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment},
  author = {Lancheng Gao and Ziheng Jia and Zixuan Xing and Wei Sun and Huiyu Duan and Guangtao Zhai and Xiongkuo Min},
  journal= {arXiv preprint arXiv:2602.01173},
  year   = {2026}
}