English

AesTest: Measuring Aesthetic Intelligence from Perception to Production

Computer Vision and Pattern Recognition 2025-11-11 v1

Abstract

Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aesthetic assessment (IAA) are narrow in perception scope or lack the diversity needed to evaluate systematic aesthetic production. To address this gap, we introduce AesTest, a comprehensive benchmark for multimodal aesthetic perception and production, distinguished by the following features: 1) It consists of curated multiple-choice questions spanning ten tasks, covering perception, appreciation, creation, and photography. These tasks are grounded in psychological theories of generative learning. 2) It integrates data from diverse sources, including professional editing workflows, photographic composition tutorials, and crowdsourced preferences. It ensures coverage of both expert-level principles and real-world variation. 3) It supports various aesthetic query types, such as attribute-based analysis, emotional resonance, compositional choice, and stylistic reasoning. We evaluate both instruction-tuned IAA MLLMs and general MLLMs on AesTest, revealing significant challenges in building aesthetic intelligence. We will publicly release AesTest to support future research in this area.

Keywords

Cite

@article{arxiv.2511.06360,
  title  = {AesTest: Measuring Aesthetic Intelligence from Perception to Production},
  author = {Guolong Wang and Heng Huang and Zhiqiang Zhang and Wentian Li and Feilong Ma and Xin Jin},
  journal= {arXiv preprint arXiv:2511.06360},
  year   = {2025}
}

Comments

10 pages, 9 figures

R2 v1 2026-07-01T07:28:16.634Z