中文
相关论文

相关论文: Architecture inside the mirage: evaluating generat…

200 篇论文

The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of synthetic visual content. While these models enable innovation across industries, they…

The rapid progress and widespread availability of text-to-image (T2I) generative models have heightened concerns about the misuse of AI-generated visuals, particularly in the context of misinformation campaigns. Existing AI-generated image…

Text-to-image (T2I) models are increasingly popular, producing a large share of AI-generated images online. To compare model quality, voting-based leaderboards have become the standard, relying on anonymized model outputs for fairness. In…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Ali Naseh , Yuefeng Peng , Anshuman Suri , Harsh Chaudhari , Alina Oprea , Amir Houmansadr

Deep generative models have achieved promising results in image generation, and various generative model hubs, e.g., Hugging Face and Civitai, have been developed that enable model developers to upload models and users to download models.…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Zhi Zhou , Lan-Zhe Guo , Peng-Xiao Song , Yu-Feng Li

The unprecedented photorealistic results achieved by recent text-to-image generative systems and their increasing use as plug-and-play content creation solutions make it crucial to understand their potential biases. In this work, we…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Melissa Hall , Candace Ross , Adina Williams , Nicolas Carion , Michal Drozdzal , Adriana Romero Soriano

Image Quality Assessment (IQA) models are employed in many practical image and video processing pipelines to reduce storage, minimize transmission costs, and improve the Quality of Experience (QoE) of millions of viewers. These models are…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Krishna Srikar Durbha , Asvin Kumar Venkataramanan , Rajesh Sureddi , Alan C. Bovik

In recent years, advancements in generative artificial intelligence have led to the development of sophisticated tools capable of mimicking diverse artistic styles, opening new possibilities for digital creativity and artistic expression.…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Andrea Asperti , Franky George , Tiberio Marras , Razvan Ciprian Stricescu , Fabio Zanotti

Although recent text-to-image generative models have achieved impressive performance, they still often struggle with capturing the compositional complexities of prompts including attribute binding, and spatial relationships between…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Seyed Mohammad Hadi Hosseini , Amir Mohammad Izadi , Ali Abdollahi , Armin Saghafian , Mahdieh Soleymani Baghshah

The rapid advancement of generators (e.g., StyleGAN, Midjourney, DALL-E) has produced highly realistic synthetic images, posing significant challenges to digital media authenticity. These generators are typically based on a few core…

机器学习 · 计算机科学 2025-11-26 Hong-Hanh Nguyen-Le , Van-Tuan Tran , Dinh-Thuc Nguyen , Nhien-An Le-Khac

A common and controversial use of text-to-image models is to generate pictures by explicitly naming artists, such as "in the style of Greg Rutkowski". We introduce a benchmark for prompted-artist recognition: predicting which artist names…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Grace Su , Sheng-Yu Wang , Aaron Hertzmann , Eli Shechtman , Jun-Yan Zhu , Richard Zhang

The AI community has embraced multi-sensory or multi-modal approaches to advance this generation of AI models to resemble expected intelligent understanding. Combining language and imagery represents a familiar method for specific tasks…

计算与语言 · 计算机科学 2023-04-06 David Noever , Samantha Elizabeth Miller Noever

Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, and attributes. Existing inference-time strategies, such as…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Shantanu Jaiswal , Mihir Prabhudesai , Nikash Bhardwaj , Zheyang Qin , Amir Zadeh , Chuan Li , Katerina Fragkiadaki , Deepak Pathak

The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users. Despite the advancements of T2I models, a common issue encountered by users is the need for…

计算与语言 · 计算机科学 2023-10-31 Wanrong Zhu , Xinyi Wang , Yujie Lu , Tsu-Jui Fu , Xin Eric Wang , Miguel Eckstein , William Yang Wang

The comparative study of generative models often requires significant computational resources, creating a barrier for researchers and practitioners. This paper introduces GANji, a lightweight framework for benchmarking foundational AI image…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Chandon Hamel , Mike Busch

Recent advancements in multimodal Generative AI have the potential to democratize specialized architectural tasks, such as interpreting technical drawings and creating 3D CAD models, which traditionally require expert knowledge. This paper…

人工智能 · 计算机科学 2025-03-05 Jingfei Huang , Alexandros Haridis

With the rapid development of generative models, discerning AI-generated content has evoked increasing attention from both industry and academia. In this paper, we conduct a sanity check on "whether the task of AI-generated image detection…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Shilin Yan , Ouxiang Li , Jiayin Cai , Yanbin Hao , Xiaolong Jiang , Yao Hu , Weidi Xie

Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in training data. We introduce T2I-BiasBench, a unified evaluation framework of thirteen…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Nihal Jaiswal , Siddhartha Arjaria , Gyanendra Chaubey , Ankush Kumar , Aditya Singh , Anchal Chaurasiya

Generative AI models have shown impressive ability to produce images with text prompts, which could benefit creativity in visual art creation and self-expression. However, it is unclear how precisely the generated images express contexts…

人机交互 · 计算机科学 2023-03-21 Yunlong Wang , Shuyuan Shen , Brian Y. Lim

Generative AI (GenAI) creates full content based on compact prompts. While GenAI has been used for applications where the generated content is returned to the prompt sender, it can play a vital role in extending the capacity of…

网络与互联网体系结构 · 计算机科学 2026-03-13 Mathias Thorsager , Israel Leyva-Mayorga , Petar Popovski

GANs (Generative adversarial networks) is a new AI technology that can perform deep learning with less training data and has the capability of achieving transformation between two image sets. Using GAN we have carried out a comparison…

计算机视觉与模式识别 · 计算机科学 2020-05-06 Mai Cong Hung , Ryohei Nakatsu , Naoko Tosa , Takashi Kusumi , Koji Koyamada