English
Related papers

Related papers: Architecture inside the mirage: evaluating generat…

200 papers

Recent advances in text-to-image generators have led to substantial capabilities in image generation. However, the complexity of prompts acts as a bottleneck in the quality of images generated. A particular under-explored facet is the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Tham Yik Foong , Shashank Kotyan , Po Yuan Mao , Danilo Vasconcellos Vargas

From a simple text prompt, generative-AI image models can create stunningly realistic and creative images bounded, it seems, by only our imagination. These models have achieved this remarkable feat thanks, in part, to the ingestion of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Matyas Bohacek , Hany Farid

Text-to-image generation models that generate images based on prompt descriptions have attracted an increasing amount of attention during the past few months. Despite their encouraging performance, these models raise concerns about the…

Cryptography and Security · Computer Science 2023-01-10 Zeyang Sha , Zheng Li , Ning Yu , Yang Zhang

While generative AI systems have gained popularity in diverse applications, their potential to produce harmful outputs limits their trustworthiness and utility. A small but growing line of research has explored tools and processes to better…

Human-Computer Interaction · Computer Science 2025-11-27 Matheus Kunzler Maldaner , Wesley Hanwen Deng , Jason I. Hong , Kenneth Holstein , Motahhare Eslami

The rapid advancement of Text-to-Image(T2I) generative models has enabled the synthesis of high-quality images guided by textual descriptions. Despite this significant progress, these models are often susceptible in generating contents that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Yichen Sun , Zhixuan Chu , Zhan Qin , Kui Ren

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Zehan Wang , Jiayang Xu , Ziang Zhang , Tianyu Pang , Chao Du , Hengshuang Zhao , Zhou Zhao

With AI-generated content becoming ubiquitous across the web, social media, and other digital platforms, it is vital to examine how such content are inspired and generated. The creation of AI-generated images often involves refining the…

Artificial Intelligence · Computer Science 2025-04-30 Khoi Trinh , Scott Seidenberger , Raveen Wijewickrama , Murtuza Jadliwala , Anindya Maiti

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yixin Wan , Kai-Wei Chang

Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models still struggle with prompts that require rich world knowledge and implicit reasoning: both of which are critical for producing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Daoan Zhang , Che Jiang , Ruoshi Xu , Biaoxiang Chen , Zijian Jin , Yutian Lu , Jianguo Zhang , Liang Yong , Jiebo Luo , Shengda Luo

The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to bridge understanding and generation tasks. Recent advances,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Sixiang Chen , Jinbin Bai , Zhuoran Zhao , Tian Ye , Qingyu Shi , Donghao Zhou , Wenhao Chai , Xin Lin , Jianzong Wu , Chao Tang , Shilin Xu , Tao Zhang , Haobo Yuan , Yikang Zhou , Wei Chow , Linfeng Li , Xiangtai Li , Lei Zhu , Lu Qi

High-stakes decision systems increasingly require structured justification, traceability, and auditability to ensure accountability and regulatory compliance. Formal arguments commonly used in the certification of safety-critical systems…

Artificial Intelligence · Computer Science 2026-04-07 Mahyar T. Moghaddam

While recent research suggests Large Language Models match human creative performance in divergent thinking tasks, visual creativity remains underexplored. This study compared image generation in human participants (Visual Artists and Non…

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Jingjing Chang , Yixiao Fang , Peng Xing , Shuhan Wu , Wei Cheng , Rui Wang , Xianfang Zeng , Gang Yu , Hai-Bao Chen

This paper presents SCHEMA (Structured Components for Harmonized Engineered Modular Architecture), a structured prompt engineering methodology specifically developed for Google Gemini 3 Pro Image. Unlike generic prompt guidelines or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Luca Cazzaniga

The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse. In particular the generation of credible misinformation in the form of images poses a significant threat to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Sina Mavali , Jonas Ricker , David Pape , Asja Fischer , Lea Schönherr

Despite significant progress in generative AI, comprehensive evaluation remains challenging because of the lack of effective metrics and standardized benchmarks. For instance, the widely-used CLIPScore measures the alignment between a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Zhiqiu Lin , Deepak Pathak , Baiqi Li , Jiayao Li , Xide Xia , Graham Neubig , Pengchuan Zhang , Deva Ramanan

In the creative practice of text-to-image (TTI) generation, images are synthesized from textual prompts. By design, TTI models always yield an output, even if the prompt contains unknown terms. In this case, the model may generate default…

Human-Computer Interaction · Computer Science 2026-01-27 Hannu Simonen , Atte Kiviniemi , Hannah Johnston , Helena Barranha , Jonas Oppenlaender

Generative AI is reshaping how computing systems are designed, optimized, and built, yet research remains fragmented across software, architecture, and chip design communities. This paper takes a cross-stack perspective, examining how…

The extraordinary ability of generative models to generate photographic images has intensified concerns about the spread of disinformation, thereby leading to the demand for detectors capable of distinguishing between AI-generated fake…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Mingjian Zhu , Hanting Chen , Qiangyu Yan , Xudong Huang , Guanyu Lin , Wei Li , Zhijun Tu , Hailin Hu , Jie Hu , Yunhe Wang

We performed a billion locality sensitive hash comparisons between artificially generated data samples to answer the critical question - can we reproduce the results of generative AI models? Reproducibility is one of the pillars of…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-02-07 Edward Kim , Isamu Isozaki , Naomi Sirkin , Michael Robson