English

Evaluating Attribute Confusion in Fashion Text-to-Image Generation

Computer Vision and Pattern Recognition 2025-07-10 v1

Abstract

Despite the rapid advances in Text-to-Image (T2I) generation models, their evaluation remains challenging in domains like fashion, involving complex compositional generation. Recent automated T2I evaluation methods leverage pre-trained vision-language models to measure cross-modal alignment. However, our preliminary study reveals that they are still limited in assessing rich entity-attribute semantics, facing challenges in attribute confusion, i.e., when attributes are correctly depicted but associated to the wrong entities. To address this, we build on a Visual Question Answering (VQA) localization strategy targeting one single entity at a time across both visual and textual modalities. We propose a localized human evaluation protocol and introduce a novel automatic metric, Localized VQAScore (L-VQAScore), that combines visual localization with VQA probing both correct (reflection) and miss-localized (leakage) attribute generation. On a newly curated dataset featuring challenging compositional alignment scenarios, L-VQAScore outperforms state-of-the-art T2I evaluation methods in terms of correlation with human judgments, demonstrating its strength in capturing fine-grained entity-attribute associations. We believe L-VQAScore can be a reliable and scalable alternative to subjective evaluations.

Keywords

Cite

@article{arxiv.2507.07079,
  title  = {Evaluating Attribute Confusion in Fashion Text-to-Image Generation},
  author = {Ziyue Liu and Federico Girella and Yiming Wang and Davide Talon},
  journal= {arXiv preprint arXiv:2507.07079},
  year   = {2025}
}

Comments

Accepted to ICIAP25. Project page: site [https://intelligolabs.github.io/L-VQAScore/\

R2 v1 2026-07-01T03:53:36.628Z