English

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

Computation and Language 2025-05-22 v2 Artificial Intelligence

Abstract

Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demonstrated strong reasoning abilities in STEM tasks (e.g. mathematics and coding), they still struggle to generate creative expressions such as resonant jokes and insightful satire. Moreover, existing benchmarks are constrained by their limited modalities and insufficient categories, hindering the exploration of comprehensive creativity in video-based Comment Art creation. To address these limitations, we introduce GODBench, a novel benchmark that integrates video and text modalities to systematically evaluate MLLMs' abilities to compose Comment Art. Furthermore, inspired by the propagation patterns of waves in physics, we propose Ripple of Thought (RoT), a multi-step reasoning framework designed to enhance the creativity of MLLMs. Extensive experiments reveal that existing MLLMs and CoT methods still face significant challenges in understanding and generating creative video comments. In contrast, RoT provides an effective approach to improve creative composing, highlighting its potential to drive meaningful advancements in MLLM-based creativity. GODBench is publicly available at https://github.com/stan-lei/GODBench-ACL2025.

Keywords

Cite

@article{arxiv.2505.11436,
  title  = {GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art},
  author = {Yiming Lei and Chenkai Zhang and Zeming Liu and Haitao Leng and Shaoguo Liu and Tingting Gao and Qingjie Liu and Yunhong Wang},
  journal= {arXiv preprint arXiv:2505.11436},
  year   = {2025}
}

Comments

69 pages, 66 figures, accepted by ACL 2025

R2 v1 2026-06-28T23:36:22.761Z