English
Related papers

Related papers: Re-imagine the Negative Prompt Algorithm: Transfor…

200 papers

Negative Prompting (NP) is widely utilized in diffusion models, particularly in text-to-image applications, to prevent the generation of undesired features. In this paper, we show that conventional NP is limited by the assumption of a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Felix Koulischer , Johannes Deleu , Gabriel Raya , Thomas Demeester , Luca Ambrogioni

Well-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation,…

Computation and Language · Computer Science 2024-01-01 Yaru Hao , Zewen Chi , Li Dong , Furu Wei

While large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generating unsafe content, violating copyright, and perpetuating…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Ruchika Chavhan , Da Li , Timothy Hospedales

In text-to-image generation, different initial noises induce distinct denoising paths with a pretrained Stable Diffusion (SD) model. While this pattern could output diverse images, some of them may fail to align well with the prompt.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yunze Tong , Didi Zhu , Zijing Hu , Jinluan Yang , Ziyu Zhao

Text-to-image diffusion models excel at generating high-quality, diverse images from natural language prompts. However, they often fail to produce semantically accurate results when the prompt contains concept combinations that contradict…

Graphics · Computer Science 2026-03-25 Saar Huberman , Or Patashnik , Omer Dahary , Ron Mokady , Daniel Cohen-Or

Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appearance, reference images are usually leveraged for the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Zhiqi Huang , Huixin Xiong , Haoyu Wang , Longguang Wang , Zhiheng Li

For text-to-image generation, automatically refining user-provided natural language prompts into the keyword-enriched prompts favored by systems is essential for the user experience. Such a prompt refinement process is analogous to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jingtao Zhan , Qingyao Ai , Yiqun Liu , Yingwei Pan , Ting Yao , Jiaxin Mao , Shaoping Ma , Tao Mei

Text-to-3D synthesis has recently emerged as a new approach to sampling 3D models by adopting pretrained text-to-image models as guiding visual priors. An intriguing but underexplored problem with existing text-to-3D methods is that 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Uy Dieu Tran , Minh Luu , Phong Ha Nguyen , Khoi Nguyen , Binh-Son Hua

Diffusion models, which have emerged to become popular text-to-image generation models, can produce high-quality and content-rich images guided by textual prompts. However, there are limitations to semantic understanding and commonsense…

Computation and Language · Computer Science 2023-11-30 Shanshan Zhong , Zhongzhan Huang , Wushao Wen , Jinghui Qin , Liang Lin

Many deep learning tasks require annotations that are too time consuming for human operators, resulting in small dataset sizes. This is especially true for dense regression problems such as crowd counting which requires the location of…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Arian Bakhtiarnia , Qi Zhang , Alexandros Iosifidis

In the evolving landscape of text-to-3D technology, Dreamfusion has showcased its proficiency by utilizing Score Distillation Sampling (SDS) to optimize implicit representations such as NeRF. This process is achieved through the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Yuzhong Huang , Zhong Li , Zhang Chen , Zhiyuan Ren , Guosheng Lin , Fred Morstatter , Yi Xu

Leveraging multi-view diffusion models as priors for 3D optimization have alleviated the problem of 3D consistency, e.g., the Janus face problem or the content drift problem, in zero-shot text-to-3D models. However, the 3D geometric…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Seungwook Kim , Kejie Li , Xueqing Deng , Yichun Shi , Minsu Cho , Peng Wang

Existing neural rendering-based text-to-3D-portrait generation methods typically make use of human geometry prior and diffusion models to obtain guidance. However, relying solely on geometry information introduces issues such as the Janus…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Yiqian Wu , Hao Xu , Xiangjun Tang , Xien Chen , Siyu Tang , Zhebin Zhang , Chen Li , Xiaogang Jin

The advancements in the domain of LLMs in recent years have surprised many, showcasing their remarkable capabilities and diverse applications. Their potential applications in various real-world scenarios have led to significant research on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Sujith Vemishetty , Advitiya Arora , Anupama Sharma

Text-driven 3D scene generation is widely applicable to video gaming, film industry, and metaverse applications that have a large demand for 3D scenes. However, existing text-to-3D generation methods are limited to producing 3D objects with…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Jingbo Zhang , Xiaoyu Li , Ziyu Wan , Can Wang , Jing Liao

Diffusion models (DMs) have demonstrated an unparalleled ability to create diverse and high-fidelity images from text prompts. However, they are also well-known to vary substantially regarding both prompt adherence and quality. Negative…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Alakh Desai , Nuno Vasconcelos

Decompositional reconstruction of 3D scenes, with complete shapes and detailed texture of all objects within, is intriguing for downstream applications but remains challenging, particularly with sparse views as input. Recent approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Junfeng Ni , Yu Liu , Ruijie Lu , Zirui Zhou , Song-Chun Zhu , Yixin Chen , Siyuan Huang

Denoising in the sRGB image space is challenging due to large noise variability. Although end-to-end methods perform well, their effectiveness in real-world scenarios is limited by the scarcity of real noisy-clean image pairs, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jaekyun Ko , Dongjin Kim , Soomin Lee , Guanghui Wang , Tae Hyun Kim

Text-to-image diffusion models have achieved remarkable performance in image synthesis, while the text interface does not always provide fine-grained control over certain image factors. For instance, changing a single token in the text can…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Chen Wu , Fernando De la Torre

Recently, there has been a significant advancement in text-to-image diffusion models, leading to groundbreaking performance in 2D image generation. These advancements have been extended to 3D models, enabling the generation of novel 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Jangho Park , Gihyun Kwon , Jong Chul Ye