English
Related papers

Related papers: A Creative Agent is Worth a 64-Token Template

200 papers

Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as synthetic vision data generators? In this work, we revisit the promise of synthetic data as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Krzysztof Adamkiewicz , Brian Bernhard Moser , Stanislav Frolov , Tobias Christian Nauen , Federico Raue , Andreas Dengel

Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Chuancheng Shi , Shangze Li , Shiming Guo , Simiao Xie , Wenhua Wu , Jingtong Dou , Chao Wu , Canran Xiao , Cong Wang , Zifeng Cheng , Fei Shen , Tat-Seng Chua

Effective prompt design is essential for improving the planning capabilities of large language model (LLM)-driven agents. However, existing structured prompting strategies are typically limited to single-agent, plan-only settings, and often…

Artificial Intelligence · Computer Science 2025-07-08 Bruce Yang , Xinfeng He , Huan Gao , Yifan Cao , Xiaofan Li , David Hsu

Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lacking compositionality could have severe implications for…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Zhiheng Li , Martin Renqiang Min , Kai Li , Chenliang Xu

Concept-based Explainable Artificial Intelligence (XAI) interprets deep learning models using human-understandable visual features (e.g., textures or object parts) by linking internal representations to class predictions, thereby bridging…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Giacomo Astolfi , Matteo Bianchi , Riccardo Campi , Antonio De Santis , Marco Brambilla

The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users. Despite the advancements of T2I models, a common issue encountered by users is the need for…

Computation and Language · Computer Science 2023-10-31 Wanrong Zhu , Xinyi Wang , Yujie Lu , Tsu-Jui Fu , Xin Eric Wang , Miguel Eckstein , William Yang Wang

Blending visual and textual concepts into a new visual concept is a unique and powerful trait of human beings that can fuel creativity. However, in practice, cross-modal conceptual blending for humans is prone to cognitive biases, like…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Wonwoong Cho , Yanxia Zhang , Yan-Ying Chen , David I. Inouye

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Tao Liu , Kai Wang , Senmao Li , Joost van de Weijer , Fahad Shahbaz Khan , Shiqi Yang , Yaxing Wang , Jian Yang , Ming-Ming Cheng

Text-to-Image (T2I) models have made significant advancements in recent years, but they still struggle to accurately capture intricate details specified in complex compositional prompts. While fine-tuning T2I models with reward objectives…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Luca Eyring , Shyamgopal Karthik , Karsten Roth , Alexey Dosovitskiy , Zeynep Akata

Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Zhipeng Chen , Lan Yang , Yonggang Qi , Honggang Zhang , Kaiyue Pang , Ke Li , Yi-Zhe Song

A significant ``modality gap" exists between the abundance of text-only data and the increasing power of multimodal models. This work systematically investigates whether images generated on-the-fly by Text-to-Image (T2I) models can serve as…

Multimedia · Computer Science 2026-03-04 Yuesheng Huang , Peng Zhang , Xiaoxin Wu , Riliang Liu , Jiaqi Liang

This work addresses the challenge of quantifying originality in text-to-image (T2I) generative diffusion models, with a focus on copyright originality. We begin by evaluating T2I models' ability to innovate and generalize through controlled…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Adi Haviv , Shahar Sarfaty , Uri Hacohen , Niva Elkin-Koren , Roi Livni , Amit H Bermano

Recent years have witnessed the substantial progress of large-scale models across various domains, such as natural language processing and computer vision, facilitating the expression of concrete concepts. Unlike concrete concepts that are…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Jiayi Liao , Xu Chen , Qiang Fu , Lun Du , Xiangnan He , Xiang Wang , Shi Han , Dongmei Zhang

Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Sedat Porikli , Vedat Porikli

We present a novel task and benchmark for evaluating the ability of text-to-image(T2I) generation models to produce images that align with commonsense in real life, which we call Commonsense-T2I. Given two adversarial text prompts…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Xingyu Fu , Muyu He , Yujie Lu , William Yang Wang , Dan Roth

Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Anne-Sofie Maerten , Juliane Verwiebe , Shyamgopal Karthik , Ameya Prabhu , Johan Wagemans , Matthias Bethge

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applications in world…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Fanqing Meng , Wenqi Shao , Lixin Luo , Yahong Wang , Yiran Chen , Quanfeng Lu , Yue Yang , Tianshuo Yang , Kaipeng Zhang , Yu Qiao , Ping Luo

Beyond conveying semantic information, images also possess cognitive properties that elicit specific psychological responses from viewers, such as memory encoding or emotional reactions. Although modern text-to-image (T2I) models generate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Shengqi Dang , Yi He , Jiaying Lei , Ziqing Qian , Nan Cao

Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, image quality, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Ziwei Huang , Wanggui He , Quanyu Long , Yandi Wang , Haoyuan Li , Zhelun Yu , Fangxun Shu , Long Chan , Hao Jiang , Fei Wu , Leilei Gan

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Lifeng Chen , Jiner Wang , Zihao Pan , Beier Zhu , Xiaofeng Yang , Chi Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›