English

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Artificial Intelligence 2025-11-18 v1

Abstract

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists of two key components: 1) an evaluation benchmark covering the multiple dimensions from creative idea to process to products; 2) CreMIT (Creativity Multimodal Instruction Tuning dataset), a multimodal creativity evaluation dataset, consisting of 2.2K diverse-sourced multimodal data, 79.2K human feedbacks and 4.7M multi-typed instructions. Specifically, to ensure MLLMs can handle diverse creativity-related queries, we prompt GPT to refine these human feedbacks to activate stronger creativity assessment capabilities. CreBench serves as a foundation for building MLLMs that understand human-aligned creativity. Based on the CreBench, we fine-tune open-source general MLLMs, resulting in CreExpert, a multimodal creativity evaluation expert model. Extensive experiments demonstrate that the proposed CreExpert models achieve significantly better alignment with human creativity evaluation compared to state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision.

Keywords

Cite

@article{arxiv.2511.13626,
  title  = {CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product},
  author = {Kaiwen Xue and Chenglong Li and Zhonghong Ou and Guoxin Zhang and Kaoyan Lu and Shuai Lyu and Yifan Zhu and Ping Zong Junpeng Ding and Xinyu Liu and Qunlin Chen and Weiwei Qin and Yiran Shen and Jiayi Cen},
  journal= {arXiv preprint arXiv:2511.13626},
  year   = {2025}
}

Comments

13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

R2 v1 2026-07-01T07:41:39.145Z