中文
相关论文

相关论文: ContextGen: Contextual Layout Anchoring for Identi…

200 篇论文

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Abdelrahman Eldesokey , Peter Wonka

Large Language models (LLMs) have achieved encouraging results in tabular data generation. However, existing approaches require fine-tuning, which is computationally expensive. This paper explores an alternative: prompting a fixed LLM with…

机器学习 · 计算机科学 2025-02-25 Liancheng Fang , Aiwei Liu , Hengrui Zhang , Henry Peng Zou , Weizhi Zhang , Philip S. Yu

Existing approaches for controlling text-to-image diffusion models, while powerful, do not allow for explicit 3D object-centric control, such as precise control of object orientation. In this work, we address the problem of multi-object…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Rishubh Parihar , Vaibhav Agrawal , Sachidanand VS , R. Venkatesh Babu

Recent text-to-image diffusion models are able to learn and synthesize images containing novel, personalized concepts (e.g., their own pets or specific items) with just a few examples for training. This paper tackles two interconnected…

计算机视觉与模式识别 · 计算机科学 2024-02-26 Chun-Hsiao Yeh , Ta-Ying Cheng , He-Yen Hsieh , Chuan-En Lin , Yi Ma , Andrew Markham , Niki Trigoni , H. T. Kung , Yubei Chen

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and often degrade image quality, and (ii) current techniques…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Ruidong Chen , Yancheng Bai , Xuanpu Zhang , Jianhao Zeng , Lanjun Wang , Dan Song , Lei Sun , Xiangxiang Chu , Anan Liu

This research focuses on the development and enhancement of text-to-image denoising diffusion models, addressing key challenges such as limited sample diversity and training instability. By incorporating Classifier-Free Guidance (CFG) and…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Rajdeep Roshan Sahu

In computer vision, it is well-known that a lack of data diversity will impair model performance. In this study, we address the challenges of enhancing the dataset diversity problem in order to benefit various downstream tasks such as…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Yuhang Li , Xin Dong , Chen Chen , Weiming Zhuang , Lingjuan Lyu

We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images…

计算机视觉与模式识别 · 计算机科学 2022-06-03 Zuopeng Yang , Daqing Liu , Chaoyue Wang , Jie Yang , Dacheng Tao

Augmentation for dense prediction typically relies on either sample mixing or generative synthesis. Mixing improves robustness but misaligned masks yield soft label ambiguity. Diffusion synthesis increases apparent diversity but, when…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Pengyu Jie , Wanquan Liu , Rui He , Yihui Wen , Deyu Meng , Chenqiang Gao

Addressing the limitations of text as a source of accurate layout representation in text-conditional diffusion models, many works incorporate additional signals to condition certain attributes within a generated image. Although successful,…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Jonghyun Lee , Hansam Cho , Youngjoon Yoo , Seoung Bum Kim , Yonghyun Jeong

Text-to-image diffusion models have recently taken center stage as pivotal tools in promoting visual creativity across an array of domains such as comic book artistry, children's literature, game development, and web design. These models…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Kiymet Akdemir , Pinar Yanardag

Despite recent progress in diffusion models, generating realistic head portraits from novel viewpoints remains a significant challenge. Most current approaches are constrained to limited angular ranges, predominantly focusing on frontal or…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Stathis Galanakis , Alexandros Lattas , Stylianos Moschoglou , Bernhard Kainz , Stefanos Zafeiriou

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diffusion-based models that treat layout planning and…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Runze He , Bo Cheng , Yuhang Ma , Qingxiang Jia , Shanyuan Liu , Ao Ma , Xiaoyu Wu , Liebucha Wu , Dawei Leng , Yuhui Yin

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Minjung Shin , Hyunin Cho , Sooyeon Go , Jin-Hwa Kim , Youngjung Uh

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Hongji Yang , Yucheng Zhou , Wencheng Han , Runzhou Tao , Zhongying Qiu , Jianfei Yang , Jianbing Shen

Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation using diffusion models, aiming to streamline creative…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Hui Zhang , Dexiang Hong , Maoke Yang , Yutao Cheng , Zhao Zhang , Weidong Chen , Jie Shao , Xinglong Wu , Zuxuan Wu , Yu-Gang Jiang

Interior design is a complex and creative discipline involving aesthetics, functionality, ergonomics, and materials science. Effective solutions must meet diverse requirements, typically producing multiple deliverables such as renderings…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yuxuan Yang , Tao Geng

Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual encoder; while it can capture coarse global information, it…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Vatsal Agarwal , Matthew Gwilliam , Gefen Kohavi , Eshan Verma , Daniel Ulbricht , Abhinav Shrivastava

Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Yuanzhi Zhu , Zhaohai Li , Tianwei Wang , Mengchao He , Cong Yao

In this paper, we introduce StableGarment, a unified framework to tackle garment-centric(GC) generation tasks, including GC text-to-image, controllable GC text-to-image, stylized GC text-to-image, and robust virtual try-on. The main…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Rui Wang , Hailong Guo , Jiaming Liu , Huaxia Li , Haibo Zhao , Xu Tang , Yao Hu , Hao Tang , Peipei Li